Optical module junction temperature cloud coordination and regulation system and method

By using a cloud-based collaborative control system for optical module junction temperature, the system monitors and generates control signals in real time using cloud analysis, thus solving the problem of response lag in the optical module heat dissipation system under dynamic loads and achieving stable control of junction temperature and improved energy efficiency.

CN121680529BActive Publication Date: 2026-05-12ZHEJIANG XINHAN INTELLIGENT TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG XINHAN INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2026-02-10
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing optical module heat dissipation systems cannot respond quickly to dynamic load changes, causing chip junction temperatures to exceed safe thresholds, affecting performance and lifespan, and lacking adaptive adjustment capabilities.

Method used

A temperature monitoring module collects junction temperature data in real time, uploads it to a cloud analysis module via a wireless network for machine learning analysis, generates control signals, and adjusts the state of heat dissipation devices through a local execution module to achieve prediction and dynamic control of the optical module junction temperature.

Benefits of technology

Stable control of the junction temperature of the optical module has been achieved, which has improved the long-term reliability and performance stability, reduced the risk of equipment downtime, and improved overall energy efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121680529B_ABST
    Figure CN121680529B_ABST
Patent Text Reader

Abstract

The application relates to a light module junction temperature cloud coordination and control system and method, in particular to the field of light module heat dissipation, reliable temperature data is obtained through local real-time sensing and signal conditioning, and efficient and reliable coordination with the cloud is realized by relying on a wireless network, historical data and intelligent algorithms are aggregated in the cloud, a forward-looking optimization control strategy is generated, and the local execution unit is fed down, finally, the working state of the heat dissipation device is accurately adjusted, so that the junction temperature can be continuously and stably maintained in a safe range when facing dynamic load and environmental changes, thereby significantly improving the long-term working reliability, performance stability and overall energy efficiency of the light module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of optical module heat dissipation, and more specifically, to a cloud-based collaborative control system and method for optical module junction temperature. Background Technology

[0002] With the rapid development of data centers, high-performance computing, and 5G / 6G communication networks, optical modules are evolving towards higher transmission rates, higher integration, and higher power density. For example, the power consumption of a 1.6T optical module has exceeded 40 watts, and its core chips, such as the digital signal processor and laser driver, generate a large amount of concentrated heat dissipation within a small space. To address this challenge, heat dissipation technology has evolved from traditional air cooling solutions to enhanced air cooling solutions combined with vapor chambers, and further to liquid cooling solutions. Currently, liquid cooling has become the mainstream approach for solving high-power heat dissipation, specifically including large-area cold plate solutions designed for the entire optical module cage, and more advanced solutions that directly integrate miniaturized cold plates into the optical module housing (In-...). The in-cage liquid cooling solution typically employs a microchannel flow plate design, using CNC machining or tooth-shaving technology to form a precision fin structure. A copper cover plate combined with brazing ensures structural strength and sealing, aiming to achieve efficient and compact heat dissipation. These optical modules are densely deployed in server racks or base station equipment, and their workload and data traffic are highly dynamic. They may jump from low load to full load in milliseconds due to sudden traffic surges. At the same time, external ambient temperature fluctuations and thermal interference from adjacent devices also constitute complex external thermal boundary conditions. In this highly dynamic and high-density complex thermal environment, stable control of the junction temperature of the optical module chip is the key to ensuring its performance, reliability and lifespan.

[0003] However, the temperature control capabilities of existing optical module cooling systems face severe challenges. Currently, commonly used temperature control methods primarily rely on local proportional-integral-derivative (PI-DE) control algorithms or alarm and response mechanisms based on fixed temperature thresholds. These methods are essentially passive and lagging feedback control. Faced with the millisecond-level rapid fluctuations in optical module power consumption caused by sudden changes in service load, the response speed of traditional control loops cannot match the dynamic changes in heat sources due to delays in sensing, calculation, and execution, resulting in lagging control. More fundamentally, these methods lack the collaborative analysis and learning of historical operating data of the optical module, real-time load characteristics, and the thermal inertia of the cooling system itself. Optical communication modules, whose control strategies are typically static or have fixed parameters, cannot predict thermal trends in the short term or adapt to different operating modes and environmental conditions. Therefore, under dynamic load scenarios, the chip junction temperature is prone to instantaneously exceeding the safety threshold, such as consistently exceeding the upper limit of 85 degrees Celsius. This will directly induce chip performance degradation or even failure due to overheating, leading to a sharp increase in the bit error rate of optical communication. At the same time, high temperatures will accelerate the aging of the chip and its packaging materials, significantly shortening the lifespan of the optical module by more than 30%, and in severe cases, may cause equipment shutdown, posing a serious threat to the stable operation and energy efficiency of the entire communication system. Summary of the Invention

[0004] This invention addresses the technical problems existing in the prior art by providing a cloud-based collaborative control system and method for optical module junction temperature. The system utilizes a temperature monitoring module, a data communication module, a cloud analysis module, and a local execution module to solve the problems mentioned in the background section.

[0005] The technical solution of the present invention to solve the above-mentioned technical problems is as follows: a cloud-based collaborative control system for optical module junction temperature, specifically comprising: a temperature monitoring module, a data communication module, a cloud analysis module, and a local execution module connected in sequence, wherein;

[0006] Temperature monitoring module: When the optical module is powered on, it collects the junction temperature analog signal through the temperature sensor integrated near the optical module chip, converts the junction temperature analog signal into a digital temperature value, and generates a junction temperature data packet;

[0007] Data communication module: Receives junction temperature data packets, encapsulates the junction temperature data packets into data frames through the wireless network protocol and uploads them to the cloud server, triggering the cloud data processing flow;

[0008] Cloud analytics module: Deployed on a cloud server, it calls the historical temperature database based on the received junction temperature data packets to perform trend analysis, and uses machine learning algorithms to generate and output control signals;

[0009] Local execution module: Receives control signals from the cloud, adjusts the current of the thermoelectric cooler or the fan speed through the drive circuit, changes the heat dissipation state of the optical module, and stabilizes the junction temperature within the preset threshold range;

[0010] In a preferred embodiment, the specific operation of acquiring the junction temperature analog signal through a temperature sensor integrated near the optical module chip in the temperature monitoring module is as follows:

[0011] The temperature sensor uses a surface-mount negative temperature coefficient thermistor, which is tightly mounted on the package surface of the optical module chip to directly sense the junction temperature of the chip. When the optical module is powered on, the resistance of the thermistor changes with the junction temperature. A constant current of 100 microamps is provided through a constant current source circuit, generating a voltage drop across the thermistor. This voltage drop serves as the initial junction temperature analog signal. During the acquisition process, within each set sampling time window, multiple instantaneous voltage values ​​of this initial junction temperature analog signal are continuously acquired at the corresponding time. These are called instantaneous voltage sample values. Each instantaneous voltage sample value is assigned an exponentially decaying weighting coefficient, which is calculated based on the absolute distance between its sampling sequence number and the center sequence number of the window. Each instantaneous voltage sample value is multiplied by its corresponding weighting coefficient, summed, and then divided by the sum of all weighting coefficients to calculate and output a weighted average analog voltage signal.

[0012] The temperature monitoring module also includes a signal conditioning circuit, which contains an operational amplifier to amplify the weighted averaged analog voltage signal with a gain of 2. After amplification, the resulting analog voltage signal is used as the input for analog-to-digital conversion.

[0013] In a preferred embodiment, the specific process of converting the junction temperature analog signal into a digital temperature value to generate a junction temperature data packet is as follows:

[0014] The analog voltage signal, used as the input to the analog-to-digital converter (ADC), is input to a 16-bit resolution Σ-Δ modulation ADC. The ADC uses 5.0 volts as a voltage reference. During conversion, a zero-drift voltage compensation value of 0.1 volts is first subtracted from the voltage value of the analog voltage signal used as the ADC input to obtain the compensated voltage. The compensated voltage is then divided by the 5.0 volt voltage reference to obtain a proportional value. This proportional value is multiplied by 65535 (2 to the power of 16 minus 1), and the result is rounded to obtain a raw digital code. Then, a ternary polynomial calculation process is used to map the raw digital code to a digital temperature value. This calculation process is as follows: First, four pre-stored polynomial calibration coefficients are read, representing the coefficients for the constant term, linear term, quadratic term, and cubic term, respectively. Next, the square and cube values ​​of the raw digital code are calculated. Then, the product of the constant term coefficient, the linear term coefficient, and the raw digital code, and the product of the quadratic term coefficient and the cubic term coefficient are calculated. The product of the number and the square of the original digital code, and the product of the coefficient of the cubic term and the cube of the original digital code are summed to obtain the digital temperature value. The digital temperature value is stored internally as a 16-bit unsigned integer, with the high 8 bits representing the integer part of the temperature and the low 8 bits representing the fractional part. The conversion result is then digitally filtered using a moving average filtering algorithm with a window size of 5. This algorithm calculates the arithmetic mean of the current and the previous four consecutive digital temperature values ​​as the output value of the current filter, resulting in the filtered digital temperature value. The temperature monitoring module also includes a calibration unit, which corrects the polynomial calibration coefficients in the ternary polynomial calculation process using factory-stored calibration parameters to compensate for sensor nonlinearity errors. Finally, the filtered digital temperature value is used to assemble the junction temperature data packet. The junction temperature data packet contains, in sequence, a data packet header, a timestamp field, a digital temperature value field, and a cyclic redundancy check (CRC) field.

[0015] In a preferred embodiment, the specific operation of encapsulating junction temperature data packets into data frames via a wireless network protocol in the data communication module is as follows:

[0016] The process of receiving junction temperature data packets from the temperature monitoring module and encapsulating these packets into data frames involves sequentially performing frame header construction, payload mapping, and frame assembly.

[0017] First, based on the link layer protocol specification of the current wireless network connection, a data frame control header conforming to the protocol is generated; in addition to the source address, destination address and protocol type fields, a data feature identifier field is added to the frame control header.

[0018] Next, calculate the data feature identifier:

[0019] Obtain the complete 32-bit value of the timestamp field in the junction temperature data packet and the complete 16-bit value of the digital temperature value field. Shift the 16-bit value eight bits to the least significant bit to obtain an 8-bit temperature feature value. Perform a bitwise XOR operation between the complete 32-bit value of the timestamp field and the 8-bit temperature feature value to obtain an intermediate mixed value. Use the intermediate mixed value as input to perform a lightweight hash function operation, using the FNV-1a algorithm.

[0020] Then, they are combined to form a complete data frame:

[0021] The calculated data feature identifier is filled into the data feature identifier field of the frame control header; then, according to the link layer protocol specification, the data packet header, timestamp field, digital temperature value field and cyclic redundancy check code field contained in the junction temperature data packet are concatenated in sequence to form the payload of the data frame; finally, the filled frame control header and payload are combined in sequence to form a primary data frame that conforms to the protocol format and is ready to be sent.

[0022] In a preferred embodiment, the specific process of triggering the cloud data processing flow is as follows:

[0023] Before sending the initial data frame, the communication quality of the current wireless channel is monitored in real time to obtain a signal-to-noise ratio estimate that characterizes the channel quality.

[0024] Fragmentation Decision: Define a fragmentation decision condition, which is as follows: compare the estimated signal-to-noise ratio (SNR) value with a preset SNR threshold value; if the estimated SNR value is lower than the SNR threshold value, it is determined that the channel conditions require fragmentation, dividing the primary data frame into multiple transmission frame blocks, where the length of each transmission frame block is set to a preset fixed value; if the estimated SNR value is higher than or equal to the SNR threshold value, fragmentation is not initiated, and the entire primary data frame is treated as a single transmission frame block.

[0025] Code rate calculation: For each transmission frame block to be sent, the code rate used for forward error correction coding is determined by a dynamic calculation process. This calculation process uses a signal-to-noise ratio (SNR) estimate, an SNR threshold, a preset SNR normalization parameter, and a preset base code rate as inputs. During the calculation, firstly, the natural logarithm of the ratio of the SNR estimate to the SNR normalization parameter plus one is calculated and denoted as the first logarithm. Then, the natural logarithm of the ratio of the SNR threshold to the SNR normalization parameter plus one is calculated and denoted as the second logarithm. Finally, the base code rate is multiplied by the quotient of the first and second logarithms, and the result is the dynamic code rate to be used.

[0026] Each transmission frame block is channel-coded using the determined dynamic coding rate to generate the final transmission frame block sequence, which is then sent sequentially through the wireless network interface. During the transmission process, an automatic retransmission request mechanism is implemented, and differentiated retransmission strategies are adopted for different transmission frame blocks.

[0027] The cloud server decodes, verifies, and reassembles the transmission frame blocks at the receiving end to recover the original primary data frame. After successful reassembly, the cloud server extracts the junction temperature data packet from the frame and verifies it. Subsequently, it executes the trigger signal generation logic to determine whether to start the cloud data processing flow: First, it determines the total number of transmission frame blocks required to constitute the junction temperature data packet. Next, it assigns a weight coefficient to each transmission frame block, with the following rules: Transmission frame blocks containing complete data of the digital temperature value field in the payload have their weight coefficient set to a first fixed value; transmission frame blocks containing only the timestamp field and cyclic redundancy check (CRC) field in the payload have their weight coefficient summed to a second fixed value, and the weight coefficient of each such transmission frame block is the second fixed value divided by the number of such transmission frame blocks. Then, it defines a binary reception status flag for each transmission frame block. If the transmission frame block is successfully received and passes verification, its flag is the first status value; otherwise, it is the second status value.

[0028] Next, the weighted success rate is calculated by multiplying the weight coefficients of all transmitted frame blocks by their corresponding reception status flags. Finally, the calculated weighted success rate is compared with a preset success reception threshold. Only when the weighted success rate is greater than or equal to the success reception threshold will the cloud server automatically generate a data processing flow trigger signal with a valid value. If the weighted success rate is less than the success reception threshold, an invalid trigger signal will be generated or no signal will be generated. The data processing flow trigger signal with a valid value is used to start the cloud data processing flow for the timestamps and digital temperature values ​​contained in the junction temperature data packet.

[0029] In a preferred embodiment, the cloud analysis module's process of performing trend analysis based on the received junction temperature data packet and accessing the historical temperature database includes two stages: data fusion and cleaning, and spatiotemporal feature extraction. Specifically:

[0030] During the data fusion and cleaning phase, the cloud analysis module receives junction temperature data packets from the data communication module and extracts the current timestamp and digital temperature value. Based on the extracted current timestamp and digital temperature value, it queries the historical temperature database to obtain the historical temperature sequence of the same optical module device within a preset historical time window. At the same time, it obtains the set of digital temperature values ​​of other optical module devices deployed in the same equipment rack and with the same topological location near the current timestamp.

[0031] Flow cytometry anomaly cleaning is performed on historical temperature sequences. This process employs a modified Grubbs test method. For each digital temperature value in the historical temperature sequence designated as a check point, the following calculation and judgment steps are performed sequentially: D1. Calculate the time-weighted moving median centered on the check point. The calculation process is as follows: Assign a weight value to each other digital temperature value in the historical temperature sequence located within a time window before and after the check point. This weight value is obtained by dividing the absolute value of the difference between the timestamp of the other digital temperature value and the timestamp of the check point by a preset time decay constant, taking the negative value, and then performing an exponential operation with the natural constant e as the base. Multiply each of the other digital temperature values ​​within the window by... D1. Sum the values ​​with their corresponding weights, then divide by the sum of all weights to obtain the time-weighted moving median. D2. Calculate a robust scaling estimate. The calculation process is as follows: First, calculate the absolute difference between each other numerical temperature value within the time window and the time-weighted moving median obtained in step 1. Then, find the median from these absolute differences. Finally, multiply the median by a constant 1.4826 to obtain the robust scaling estimate. D3. Calculate the standardized deviation of the point to be inspected. The calculation process is as follows: First, calculate the absolute difference between the numerical temperature value of the point to be inspected and the time-weighted moving median obtained in D1. Then, divide this absolute difference by the robust scaling estimate obtained in D2 to obtain the standardized deviation.

[0032] After the calculation is completed, a judgment is made: if the calculated standardized deviation value is greater than a critical value, the digital temperature value of the point to be inspected is determined to be an outlier and is removed or replaced by a smoothing function with an adjacent normal value.

[0033] After cleaning, the cleaned historical temperature sequence, the set of digital temperature values ​​from nearby devices, and the current timestamp and digital temperature value are merged to form a multidimensional input dataset.

[0034] During the spatiotemporal feature extraction stage, the following operations are performed:

[0035] First, temporal attention features are extracted from the cleaned historical temperature sequence. Specifically, for each numerical temperature value in the historical temperature sequence, the encoded vector obtained by mapping the numerical temperature value and its corresponding timestamp through an embedding function is concatenated to form a combined feature vector. To calculate the attention level of any target point to another historical point in the historical temperature sequence, an attention weight calculation process is defined: the combined feature vector of the target point is multiplied by a learnable first weight matrix to obtain the query vector; the combined feature vector of the historical point is multiplied by a learnable second weight matrix to obtain the key vector; the query vector and the key vector are concatenated sequentially into a long vector, and then this long vector is combined with a learnable weight vector. A dot product operation is performed, and the resulting scalar is nonlinearly transformed using a LeakyReLU activation function. This transformation result is then exponentially operated on with the natural constant e as the base to obtain the unnormalized attention score of the historical point to the target point. After calculating the unnormalized attention score for all historical points in the historical temperature sequence using the same method, all scores are summed. Then, the unnormalized attention score of each historical point is divided by the sum, and the quotient is the final attention weight of that historical point to the target point. After calculating the attention weights between all points in the historical temperature sequence, these weights are used to perform a weighted summation of the combined feature vectors of all historical points to generate a time feature vector reflecting the dynamic changes in the temperature of the target device itself over time.

[0036] Simultaneously, spatial graph convolution feature extraction is performed to process the set of digital temperature values ​​of neighboring devices. Specifically, each neighboring optical module device is treated as a node in the graph, and the feature of this node is its digital temperature value. An edge is established between any two nodes, and the weight of this edge is determined by the physical distance between the devices represented by the two nodes and a predefined heat dissipation coupling coefficient. Based on all nodes and edges, a graph structure is constructed, and an adjacency matrix is ​​used to represent this graph structure. The value of each element in the adjacency matrix is ​​the weight of the edge between the corresponding node pair. The sum of this adjacency matrix and an identity matrix is ​​calculated to obtain an augmented adjacency matrix with self-connections. The sum of the elements in each row of the augmented adjacency matrix is ​​calculated, and a graph is constructed with... These are diagonal matrices with diagonal elements. The negative 1 / 2 power of this diagonal matrix is ​​multiplied by the augmented adjacency matrix, and then multiplied by another negative 1 / 2 power of the diagonal matrix to obtain a normalized matrix for feature propagation. The digital temperature values ​​of all nodes are arranged into a feature matrix. The normalized matrix is ​​multiplied by the feature matrix, and the result is multiplied by a learnable weight matrix. Finally, the result is passed through a non-linear activation function to generate a new feature vector for each node, aggregating information from its neighbors. From these new feature vectors, a new feature vector representing the target optical module device is extracted as a spatial feature vector reflecting the thermal interference from surrounding devices to the target optical module device.

[0037] Finally, the temporal feature vector and the spatial feature vector are fused to obtain a high-level fused feature tensor containing spatiotemporal context information.

[0038] In a preferred embodiment, the specific process of generating and outputting control signals using machine learning algorithms includes two steps: probabilistic sequence prediction and robust control optimization. Specifically:

[0039] In the probabilistic sequence prediction step, a variational Bayesian long short-term memory network model is used to process the high-level fusion feature tensor; and then it outputs two sequences: the mean sequence of the predicted junction temperature at each future time point, and the standard deviation sequence representing the prediction uncertainty;

[0040] In the robust control optimization step, a stochastic model predictive control problem is constructed, simplifying the optical module heat dissipation system into a first-order thermodynamic model. Using the junction temperature prediction mean sequence and the standard deviation sequence representing prediction uncertainty output from the probabilistic sequence prediction step as inputs, an optimization objective function is defined. The construction and solution process of this objective function is as follows:

[0041] First, define a set of future control action sequences containing K consecutive control actions, where each action represents a control instruction to be determined;

[0042] Secondly, an optimization objective function is constructed, which finds an optimal sequence from all possible future control action sequences that minimizes the comprehensive cost. The comprehensive cost is defined as the weighted sum of the expected temperature penalty cost and the expected power consumption cost within a prediction window containing K time steps in the future.

[0043] Next, the optimization objective is defined: the optimization objective is to find an optimal sequence from all possible future control action sequences that minimizes the value of the optimization objective function;

[0044] Then, optimization is performed: a stochastic optimization algorithm is used to minimize the defined objective function and calculate the optimal sequence of future control actions.

[0045] Finally, output control instructions: extract the first control action from the optimal future control action sequence obtained from the solution, and use it as an immediate control instruction that should be issued to the local execution module immediately;

[0046] The real-time control command is encapsulated into a data packet of control signals and sent to the local execution module via the communication link.

[0047] In a preferred embodiment, the specific process of receiving control signals from the cloud in the local execution module is as follows:

[0048] The system receives control signals encapsulated in data packets from the cloud analysis module via a communication interface. First, it decrypts and verifies the integrity of the data packet; upon successful decryption, it extracts the immediate control commands. Simultaneously, the local execution module synchronously acquires the latest digital temperature value from the temperature monitoring module. Then, it performs a dynamic limiting operation at the safety boundary: based on the proximity of the digital temperature value to a preset junction temperature safety threshold, it dynamically calculates a lower limit and an upper limit for the safety execution command; these lower and upper limits constitute a dynamic safety execution range. Afterward, a limiting calculation is performed... The calculation receives three inputs: the parsed immediate control instruction, the lower limit of the safe execution instruction, and the upper limit of the safe execution instruction. The immediate control instruction is compared with both the lower and upper limits of the safe execution instruction. If it is lower than the lower limit, the output value is equal to the lower limit; if it is higher than the upper limit, the output value is equal to the upper limit; if it lies between the lower and upper limits, the output value remains unchanged, representing the immediate control instruction itself. The output of this limit calculation is the control instruction to be executed after local security verification.

[0049] In a preferred embodiment, the specific process of changing the heat dissipation state of the optical module by adjusting the current of the thermoelectric cooler or the fan speed through the drive circuit is as follows:

[0050] First, perform drive quantity mapping: based on the normalized cooling demand intensity represented by the control command to be executed, linearly convert it into a requested cooling power value in units of power; and combine it with the preset inverse characteristic model of the thermoelectric cooler or fan to calculate a preliminary drive setting value.

[0051] Next, anti-saturation dynamic compensation is performed: a model reference adaptive anti-saturation compensator is introduced; this model reference adaptive anti-saturation compensator takes the initial drive setting value as the ideal input, and calculates and outputs a final drive setting value after anti-saturation compensation in real time through an adaptive adjustment law.

[0052] When both thermoelectric coolers and fans are configured as actuators, a multi-actuator coordination strategy is implemented. This strategy includes dynamic weight allocation calculation, and the specific process is as follows:

[0053] T1. Based on the digital temperature value, the energy efficiency ratio coefficient of the thermoelectric cooler and the energy efficiency ratio coefficient of the fan at the current temperature are obtained through a preset lookup table.

[0054] T2. Calculate a dynamic allocation coefficient. The numerator of this dynamic allocation coefficient is the energy efficiency ratio coefficient of the thermoelectric cooler, and the denominator is the sum of the energy efficiency ratio coefficient of the thermoelectric cooler, the energy efficiency ratio coefficient of the fan, and a positive constant. Using this dynamic allocation coefficient, the normalized cooling demand intensity is decomposed into a cooling demand component for the thermoelectric cooler and a cooling demand component for the fan. Specifically, the cooling demand component for the thermoelectric cooler is equal to the requested cooling power value multiplied by the dynamic allocation coefficient, and the cooling demand component for the fan is equal to the requested cooling power value multiplied by one minus the difference of the dynamic allocation coefficient. And calculate their respective final drive setpoints through their respective inverse characteristic models and anti-saturation compensators.

[0055] Finally, local closed-loop fine-tuning and drive are performed: the final drive setpoint is converted into a high-precision analog current signal or pulse width modulation signal and applied to the drive circuit of the thermoelectric cooler or fan; simultaneously, local fast closed-loop fine-tuning is performed, which is based on the calculation and superposition of a fine-tuning amount, specifically: first, a temperature deviation value is calculated, which is the difference between a desired temperature derived from the cooling requirements implied by the control command to be executed and a digital temperature value; at the same time, the rate of change of the temperature deviation value over time is calculated; then, based on a high-speed, low-gain proportional-derivative control law combined with a nonlinear compensation term, the fine-tuning amount is calculated; finally, the calculated fine-tuning amount is dynamically superimposed on the final drive setpoint to form a composite control signal acting on the drive circuit;

[0056] After completing the drive and fine-tuning, the status information is encapsulated: the final drive setting value, the fine-tuning amount, the real-time acquired digital temperature value, and the local controller's operating status flag are encapsulated together into a status feedback data packet; this status feedback data packet is sent to the cloud analysis module through the communication link for online learning and adaptive optimization of the cloud model.

[0057] This application also provides a cloud-based collaborative control method for the junction temperature of an optical module, which specifically includes the following steps:

[0058] Step S1: When the optical module is powered on, a surface-mount negative temperature coefficient thermistor mounted close to the surface of the optical module chip package is used to collect the junction temperature analog signal. The resistance change of the thermistor is converted into a voltage signal through a constant current source circuit to generate an initial junction temperature analog voltage. The initial junction temperature analog voltage is amplified and filtered by a signal conditioning circuit, and then converted into a digital temperature value by an analog-to-digital converter. The digital temperature value is then packaged into a junction temperature data packet containing a data packet header, a timestamp field, a digital temperature value field, and a cyclic redundancy check code field.

[0059] Step S2: Receive junction temperature data packets through the wireless network communication interface, construct a data frame control header according to the current link layer protocol specification, encapsulate the junction temperature data packets as the payload into a complete data frame, and upload it to the cloud server through the wireless network.

[0060] Step S3: In the cloud server, parse the junction temperature data packet in the received data frame, call the time series data in the historical temperature database, use machine learning algorithm to analyze the junction temperature change trend, and generate and output control signal for adjusting heat dissipation intensity.

[0061] Step S4: Locally, the optical module receives control signals from the cloud and converts these signals into corresponding current or pulse width modulation signals through the drive circuit. This adjusts the drive current of the thermoelectric cooler or the speed of the cooling fan, thereby changing the heat dissipation state of the optical module and maintaining the junction temperature within a preset safety threshold range.

[0062] The beneficial effects of this invention are: reliable temperature data is obtained through local real-time sensing and signal conditioning, and efficient and reliable collaboration with the cloud is achieved through wireless network. The cloud aggregates historical data and intelligent algorithms to generate forward-looking optimized control strategies, which are then sent to the local execution unit. Finally, by precisely adjusting the working state of the heat dissipation device, the junction temperature can be continuously and stably maintained within a safe range when facing dynamic loads and environmental changes, thereby significantly improving the long-term working reliability, performance stability and overall energy efficiency of the optical module. Attached Figure Description

[0063] Figure 1 This is a flowchart of the method of the present invention;

[0064] Figure 2 This is a block diagram of the system structure of the present invention. Detailed Implementation

[0065] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0066] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0067] In the description of this application, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0068] Example 1

[0069] This embodiment provides, for example Figure 1 The method for cloud-based collaborative control of optical module junction temperature, as shown, specifically includes the following steps:

[0070] Step S1: When the optical module is powered on, a surface-mount negative temperature coefficient thermistor mounted close to the surface of the optical module chip package is used to collect the junction temperature analog signal. The resistance change of the thermistor is converted into a voltage signal through a constant current source circuit to generate an initial junction temperature analog voltage. The initial junction temperature analog voltage is amplified and filtered by a signal conditioning circuit, and then converted into a digital temperature value by an analog-to-digital converter. The digital temperature value is then packaged into a junction temperature data packet containing a data packet header, a timestamp field, a digital temperature value field, and a cyclic redundancy check code field.

[0071] Step S2: Receive junction temperature data packets through the wireless network communication interface, construct a data frame control header according to the current link layer protocol specification, encapsulate the junction temperature data packets as the payload into a complete data frame, and upload it to the cloud server through the wireless network.

[0072] Step S3: In the cloud server, parse the junction temperature data packet in the received data frame, call the time series data in the historical temperature database, use machine learning algorithm to analyze the junction temperature change trend, and generate and output control signal for adjusting heat dissipation intensity.

[0073] Step S4: Locally, the optical module receives control signals from the cloud and converts these signals into corresponding current or pulse width modulation signals through the drive circuit. This adjusts the drive current of the thermoelectric cooler or the speed of the cooling fan, thereby changing the heat dissipation state of the optical module and maintaining the junction temperature within a preset safety threshold range.

[0074] Example 2

[0075] This embodiment provides, for example Figure 2The optical module junction temperature cloud-based collaborative control system shown includes: a temperature monitoring module, a data communication module, a cloud analysis module, and a local execution module connected in sequence; wherein;

[0076] Temperature monitoring module: When the optical module is powered on, it collects the junction temperature analog signal through the temperature sensor integrated near the optical module chip, converts the junction temperature analog signal into a digital temperature value, and generates a junction temperature data packet;

[0077] Data communication module: Receives junction temperature data packets, encapsulates the junction temperature data packets into data frames through the wireless network protocol and uploads them to the cloud server, triggering the cloud data processing flow;

[0078] Cloud analytics module: Deployed on a cloud server, it calls the historical temperature database based on the received junction temperature data packets to perform trend analysis, and uses machine learning algorithms to generate and output control signals;

[0079] Local execution module: Receives control signals from the cloud, adjusts the current of the thermoelectric cooler or the fan speed through the drive circuit, changes the heat dissipation state of the optical module, and stabilizes the junction temperature within the preset threshold range.

[0080] In this embodiment, it is specifically necessary to explain the specific operation of the temperature monitoring module in which the junction temperature analog signal is collected by the temperature sensor integrated near the optical module chip:

[0081] The temperature sensor employs a surface-mount negative temperature coefficient (NTC) thermistor, which is tightly mounted on the package surface of the optical module chip to directly sense the junction temperature of the chip. The preferred surface-mount NTC thermistor is the 0603 package, with a thermal response time of less than 100 milliseconds, enabling rapid tracking of junction temperature changes. Tight mounting refers to the use of high thermal conductivity adhesive for bonding or soldering, ensuring a thermal resistance of less than 1°C / W to achieve efficient heat conduction from the junction temperature to the sensor. When the optical module is powered on, the thermistor's resistance changes with the junction temperature. A constant current of 100 microamps is provided by a constant current source circuit, generating a voltage drop across the thermistor. This voltage drop serves as the initial analog signal of the junction temperature. The constant current source circuit uses a low-temperature drift, high-precision reference source and an operational amplifier. The output current stability is better than ±0.5%. A current value of 100 microamps was chosen to minimize the self-heating effect of the thermistor while ensuring sufficient signal amplitude, keeping its self-heating temperature rise less than 0.1℃. The voltage value of the junction temperature analog signal is negatively correlated with the junction temperature; for every 1 degree Celsius increase in junction temperature, the voltage value decreases by 10 millivolts. This negative correlation is determined by the B-value characteristic of the thermistor, and was determined through calibration experiments. The 10 millivolt / degree Celsius change rate is the average sensitivity for the selected thermistor model within the range of -40℃ to 125℃ under a constant current of 100 microamps. Its voltage amplitude range is limited to 0 volts to 3.3 volts, corresponding to a junction temperature measurement range from -40℃ to 125℃.The 3V range matches the input range of the subsequent analog-to-digital converter. Hard limiting is achieved through clamping diodes or operational amplifier output saturation characteristics in the signal conditioning circuit to prevent overvoltage damage to subsequent circuits. The range of -40℃ to 125℃ covers the industrial-grade operating temperature requirements of the optical module. During the acquisition process, within each set sampling time window, multiple instantaneous voltage values ​​of the initial junction temperature analog signal at the corresponding moment are continuously acquired, referred to as instantaneous voltage sample values. The duration of the sampling time window is configured to 1 millisecond, and the number of instantaneous voltage sample values ​​acquired within each duration is 10. The 1-millisecond sampling window is set. To capture rapid transient changes in junction temperature, such as sudden power consumption increases due to burst data transmission, 10 points are sampled within each window, corresponding to an instantaneous sampling rate of 100 kS / s. This rate is much higher than the junction temperature change frequency, satisfying the Nyquist sampling theorem and effectively reconstructing the signal. To reduce the noise impact of data points at the edge of the sampling window and increase the contribution of data points at the center, an exponentially decaying weighting coefficient is assigned to each instantaneous voltage sample value. This weighting coefficient is calculated based on the absolute distance between its sampling sequence number (1 to 10) and the window center sequence number (e.g., set to 5). Instantaneous voltage samples closer to the window center are more likely to be sampled. The higher the value, the greater the weight; the weight decreases exponentially with increasing distance. Each instantaneous voltage sample value is multiplied by its corresponding weight coefficient, summed, and then divided by the sum of all weight coefficients to calculate a weighted average analog voltage signal. The attenuation factor, used to control the weight decay rate, is set to 0.5. This exponentially weighted average algorithm is equivalent to a digital low-pass filter, effectively suppressing high-frequency random noise and preserving the true trend of signal changes better than a simple arithmetic average. The attenuation factor of 0.5 is the result of simulation and experimental optimization, which improves noise suppression (smoothing). A good balance is achieved between signal quality and signal response speed (phase delay). The initial junction temperature analog signal is filtered before and after calculation using a low-pass filter with a cutoff frequency of 100 Hz to eliminate high-frequency noise. This low-pass filter is implemented using a first-order RC passive filter or an active filter. The cutoff frequency is set to 100 Hz to filter out power switching noise (usually greater than 10 kHz) and environmental electromagnetic interference, while retaining the useful signal of junction temperature change (usually below 10 Hz). The filtering is performed once before and once after the weighted average calculation, forming a two-stage filtering to further ensure signal quality.

[0082] The temperature monitoring module also includes a signal conditioning circuit containing an operational amplifier to amplify the weighted averaged analog voltage signal. The amplification gain is 2x to match the input range of the subsequent analog-to-digital converter (ADC). The amplified signal is then used as the input for the ADC. The operational amplifier is a low-noise, low-offset precision operational amplifier with a gain of 2x. This amplifies the pre-processed analog voltage signal (typical range 0-1.65V) to 0-3.3V, fully utilizing the dynamic range of the subsequent 16-bit ADC (0-5V) and improving conversion resolution. The acquisition interval is configurable, with a default interval of 1 millisecond, but it can be dynamically adjusted according to the optical module's operating mode, such as shortening it to 500 microseconds under high-temperature conditions. The dynamic adjustment strategy for the acquisition interval is based on a preset junction temperature threshold. For example, when the digital temperature value (obtained in subsequent steps) exceeds 85°C, the system automatically shortens the acquisition interval from the default 1 millisecond to 500 microseconds to improve the monitoring frequency and control response speed. This threshold of 85°C is based on the typical safe operating junction temperature of the optical module chip.

[0083] In addition, the temperature monitoring module integrates a self-test function. Upon power-up, it verifies the normal operation of the sensor and circuit by injecting a test signal. The test signal is a 1V DC voltage. If the value of the amplified analog voltage signal after weighted averaging and calculation deviates by more than 5%, a fault alarm is triggered. The self-test function is executed during each power-up initialization. The injected 1V test signal is loaded to the front end of the signal chain through an analog switch. The system reads the output value after passing through the entire acquisition and conditioning chain and compares it with the theoretical expected value (corresponding to a 1V input, which should be 2V after being amplified by a gain of 2). The 5% deviation threshold (i.e., the range of 1.9V to 2.1V) is a comprehensive tolerance that takes into account the operational amplifier gain error, resistor tolerance, and ADC quantization error. If the limit is exceeded, an alarm is triggered by setting the status register or pulling the external pin low.

[0084] The specific process of converting the junction temperature analog signal into a digital temperature value to generate a junction temperature data packet is as follows:

[0085] The analog voltage signal, used as the input to the analog-to-digital converter (ADC), is fed into a 16-bit resolution Σ-Δ modulation ADC. This ADC uses 5.0V as its voltage reference. During conversion, a 0.1V zero-drift compensation value is subtracted from the voltage value of the analog voltage signal used as the ADC input to obtain the compensated voltage. The Σ-Δ ADC is chosen for its high resolution and good noise immunity. The 5.0V reference source needs to have high accuracy and low temperature drift characteristics. The 0.1V zero-drift compensation value is used to eliminate potential DC offset errors in the signal conditioning circuit and the ADC itself. This 0.1V compensation value is measured by measuring the ADC output under zero-point conditions (e.g., 0V input). The code value is generated and calculated in reverse, and stored in non-volatile memory. The compensated voltage is divided by a 5.0V voltage reference to obtain a proportional value. This proportional value is multiplied by 65535 (2 to the power of 16 minus 1), and the result is rounded to obtain a raw digital code in the range of 0 to 65535. This calculation process is completed by the digital logic inside the analog-to-digital converter or by an external microcontroller, realizing a linear mapping from analog voltage to digital code. The rounding operation uses the rounding method to reduce quantization error. Then, a ternary polynomial calculation process is used to map the raw digital code into a digital temperature value. The calculation process is as follows: First, four pre-stored polynomial calibration coefficients are read, which are used for the constant term, the first term, the second term, and the third term, respectively. The coefficients of the constant term and the cubic term are -40.0, 0.01, 1.5 x 10^-6, and -2.0 x 10^-10, respectively, all in degrees Celsius. These four polynomial calibration coefficients are obtained by measuring the original digital code output by the sensor at multiple calibration temperature points (e.g., -40℃, 0℃, 25℃, 85℃, 125℃) and performing curve fitting using the least squares method. The ternary polynomial model can effectively compensate for the nonlinear characteristics of the thermistor, and the fitting error can be controlled within ±0.1℃. The coefficients are stored in the system's EEPROM or Flash. Next, the square and cube values ​​of the original digital code are calculated. Then, the coefficients of the constant term and the linear term are compared with the original... The digital temperature value is obtained by summing the product of the initial digital code, the product of the quadratic coefficient and the square of the initial digital code, and the product of the cubic coefficient and the cube of the initial digital code. The sum covers a range from -40°C to 125°C. The digital temperature value is stored internally as a 16-bit unsigned integer, with the high 8 bits representing the integer part and the low 8 bits representing the fractional part, achieving a resolution of 0.1 degrees Celsius. This 16-bit fixed-point format facilitates subsequent processing and data packet encapsulation. For example, the digital temperature value 25.6°C is stored as 0x1900 (hexadecimal), where 0x19 (decimal 25) is the integer part and 0x00 (representing 0.0) is the fractional part.7℃ is stored as 0x1901; the conversion result undergoes digital filtering using a moving average filtering algorithm with a window size of 5. This involves calculating the arithmetic mean of the current and the previous four consecutive digital temperature values ​​as the filtered output value, further improving stability and accuracy, resulting in a filtered digital temperature value. Moving average filtering effectively smooths residual random noise in ADC conversion and polynomial calculations. The window size of 5 represents a trade-off between filtering effect (smoothness) and response latency (approximately 5 sampling periods). A circular buffer is used to store historical values ​​for efficient calculation. The analog-to-digital conversion sampling rate is set to 1 kHz, but can be dynamically adjusted according to the optical module load, such as increasing it to 2 kHz in high-temperature environments; 1 kHz... The default sampling rate of / s meets the monitoring needs of most operating conditions. The dynamic adjustment logic is linked to the acquisition interval adjustment: when the acquisition interval is shortened to 500 microseconds, the analog-to-digital conversion sampling rate is simultaneously increased to 2kS / s to maintain data synchronization. The temperature monitoring module also includes a calibration unit, which corrects the polynomial calibration coefficients in the ternary polynomial calculation process using factory-stored calibration parameters to compensate for sensor nonlinearity errors. The calibration process is performed by comparing with a standard temperature source to ensure that the conversion error is less than 0.5 degrees Celsius throughout the entire measurement range. Finally, the filtered digital temperature value is used to assemble the junction temperature data packet. The junction temperature data packet sequentially includes a data packet header, a timestamp field, a digital temperature value field, and a cyclic redundancy check (CRC) field. The data packet header is fixed at 0xAA55 for frame synchronization. The timestamp field is a 32-bit unsigned integer, incremented in milliseconds by a hardware timer starting from module power-on, used for data sorting and packet loss detection. The digital temperature value field is the filtered digital temperature value in the aforementioned 16-bit format. The timestamp field is 32 bits wide, and its value starts counting from the time the temperature monitoring module powers on, with the counting unit being milliseconds. The 32-bit width supports approximately 49 days of continuous counting, meeting the long-term operation requirements of the optical module, and the millisecond resolution is sufficient to distinguish consecutive data packets. The cyclic redundancy check (CRC) field is 16 bits wide, and its value is a combination of the 32-bit binary sequence of the timestamp field and the digital temperature value carrying the filtered digital temperature value. The 16-bit binary sequence of the segment is concatenated sequentially to form a 48-bit binary sequence. This 48-bit sequence is calculated using the generator polynomial based on the CCITT standard. The binary representation of this polynomial is 1 followed by 16 bits, where the 12th, 5th, and 0th bits are 1, and the remaining bits are 0. The CRC calculation uses the standard CRC-16-CCITT algorithm (generator polynomial 0x1021). During calculation, the 48-bit data (timestamp + temperature value) is used as the dividend, and a modulo-2 division operation is performed bit by bit. The resulting 16-bit remainder is the CRC checksum, appended to the end of the data packet. The receiving end (such as the main controller) can recalculate the CRC and compare it with the received checksum to verify whether the data was corrupted during transmission, ensuring data integrity.

[0086] In this embodiment, it is specifically necessary to explain the operation of encapsulating junction temperature data packets into data frames using the wireless network protocol in the data communication module as follows:

[0087] The system receives junction temperature data packets from the temperature monitoring module. These packets contain a header, a timestamp field, a digital temperature value field, and a cyclic redundancy check (CRC) field. The total length of the junction temperature data packet is 70 bytes, with a 2-byte header (fixed at 0xAA55), a 4-byte timestamp field, a 2-byte digital temperature value field, and a 2-byte CRC field. The receiving action is performed by the buffer management unit of the data communication module to ensure complete data packet storage. The process of encapsulating the junction temperature data packet into a data frame involves sequentially performing frame header construction, payload mapping, and frame assembly. These three encapsulation steps are executed sequentially in the protocol processing unit of the data communication module. The purpose is to adapt the junction temperature data packets generated by the application layer into a link layer data frame format that can be transmitted by the underlying wireless network (such as Wi-Fi or 4G / 5G).

[0088] First, based on the link layer protocol specification of the current wireless network connection, a data frame control header conforming to the protocol is generated. In addition to the source address, destination address and protocol type fields, the frame control header also adds a data feature identifier field. If the current network is Wi-Fi (IEEE 802.11), the frame control header is generated according to its MAC frame format. The source address is the MAC address of the optical module's network interface, the destination address is the MAC address of the cloud server gateway, the protocol type field identifies the upper layer protocol as a protocol specific to this invention, and the newly added 8-bit data feature identifier field is located in the optional field area of ​​the frame control header.

[0089] Next, a data feature identifier is calculated. This calculation is performed by the identifier generation unit of the data communication module, aiming to generate a unique and representative short identifier for each junction temperature data packet. This identifier is used for fast classification and indexing in the cloud, reducing the overhead of deep parsing.

[0090] The complete 32-bit value of the timestamp field in the junction temperature data packet is obtained, and the complete 16-bit value of the digital temperature value field is obtained. The 16-bit value is shifted eight bits to the least significant bit to obtain an 8-bit temperature feature value. The complete 32-bit value of the timestamp field and the 8-bit temperature feature value are XORed to obtain an intermediate mixed value. The intermediate mixed value is used as input to perform a lightweight hash function operation. The lightweight hash function adopts the FNV-1a algorithm. The output of this operation is an 8-bit unsigned integer, which serves as the data feature identifier. The temperature feature value is obtained by right-shifting the 16-bit digital temperature value by 8 bits, which essentially extracts the integer part of the temperature value. It is then XORed with the 32-bit timestamp to integrate spatiotemporal information. The FNV-1a hash algorithm is fast and has a low collision rate, making it suitable for embedded environments. The final 8-bit identifier ranges from 0 to 255.

[0091] Then, they are combined to form a complete data frame:

[0092] The calculated data feature identifier is filled into the data feature identifier field of the frame control header. Then, according to the link layer protocol specification, the data packet header, timestamp field, digital temperature value field, and cyclic redundancy check code field contained in the junction temperature data packet are concatenated in order to form the payload of the data frame. Finally, the filled frame control header and payload are combined in order to form a primary data frame that conforms to the protocol format and is ready to be sent. The payload, i.e., the 70-byte junction temperature data packet, is encapsulated as is. After the frame is assembled, the total length of the primary data frame is, for example, (frame control header length + 70) bytes under Wi-Fi. This primary data frame is the input object for subsequent adaptive transmission processing.

[0093] The specific process for triggering the cloud data processing flow is as follows:

[0094] Before sending the initial data frame, the communication quality of the current wireless channel is monitored in real time to obtain a signal-to-noise ratio (SNR) estimate that characterizes the channel quality. This SNR estimate is expressed in decibels. The SNR estimate is measured by the physical layer chip of the wireless network interface and is updated multiple times per second, for example, 100 times per second, to reflect the rapid changes in the channel.

[0095] Fragmentation Decision: Define a fragmentation decision condition, specifically: compare the estimated signal-to-noise ratio (SNR) value with a preset SNR threshold value, expressed in decibels (dB). The preset SNR threshold value is 10 dB, determined based on extensive experiments. When the SNR is below this value, the bit error rate of transmitting the complete frame without fragmentation will significantly increase. If the estimated SNR value is lower than the SNR threshold value, it is determined that the channel conditions require fragmentation, dividing the primary data frame into multiple transmission frame blocks, each... The length of the transmission frame block is set to a preset fixed value, which is less than the length of the primary data frame when it is treated as a single frame block without fragmentation. For example, the preset fixed length value is set to 256 bytes. During fragmentation, starting from the beginning of the primary data frame, every 256 bytes (padding bytes are used to make up the difference if necessary) is divided into a transmission frame block. The last block may be less than 256 bytes. If the signal-to-noise ratio estimate is higher than or equal to the signal-to-noise ratio threshold, fragmentation is not initiated, and the entire primary data frame is treated as a single transmission frame block.

[0096] Code rate calculation: For each transmission frame block to be sent, the code rate used for forward error correction coding is determined by a dynamic calculation process. This process uses a signal-to-noise ratio (SNR) estimate, an SNR threshold, a preset SNR normalization parameter, and a preset base code rate as inputs. Both the SNR normalization parameter and the base code rate are preset constant values. The SNR normalization parameter is preset to 5 dB to adjust the function's sensitivity to SNR, and the base code rate is preset to 1 / 2, meaning a 1 / 2 code rate is used under threshold channel conditions. During calculation, first, the natural logarithm of the ratio of the SNR estimate to the SNR normalization parameter plus one is calculated and recorded as the first logarithm. Then, the SNR threshold and... The natural logarithm of the ratio of the signal-to-noise ratio (SNR) normalization parameter plus one is denoted as the second logarithm. Finally, the base coding rate is multiplied by the quotient of the first and second logarithms, and the result is the dynamic coding rate to be used. According to this calculation process, when the SNR estimate decreases, the calculated dynamic coding rate decreases, which means that the coding redundancy increases and the error correction capability is enhanced. When the SNR estimate increases, the calculated dynamic coding rate increases, which means that the coding redundancy decreases and the transmission efficiency is improved. This dynamic coding rate calculation simulates the behavior of the Shannon channel capacity formula, making the coding efficiency approach the theoretical limit of the channel. The calculated dynamic coding rate range is limited to between 1 / 4 and 3 / 4 to ensure the feasibility of the actual codec.

[0097] Each transmission frame block is channel-coded using the determined dynamic coding rate to generate the final transmission frame block sequence, which is then sent sequentially through the wireless network interface. Channel coding employs high-performance forward error correction codes such as Turbo codes or LDPC codes, and the encoder generates corresponding redundant bits based on the dynamic coding rate parameters. During transmission, an automatic retransmission request mechanism is implemented to ensure data transmission reliability, and differentiated retransmission strategies are adopted for different transmission frame blocks: transmission frame blocks containing complete data of the digital temperature value field in the payload are assigned high priority and configured with a shorter retransmission wait time; transmission frame blocks containing only timestamp and cyclic redundancy check (CRC) fields in the payload are assigned standard priority and configured with a relatively longer retransmission wait time. The retransmission wait time for high-priority frame blocks is set to 20 milliseconds, and for standard-priority frame blocks, it is set to 50 milliseconds. This strategy ensures that core temperature data can be retransmitted and recovered more quickly.

[0098] The cloud server decodes, verifies, and reassembles the transmitted frame blocks at the receiving end to recover the original primary data frame. After successful reassembly, the cloud server extracts the junction temperature data packet from the frame and verifies it. The cloud server verifies the integrity of the junction temperature data packet by checking the cyclic redundancy check (CRC) field. Subsequently, it executes the trigger signal generation logic to determine whether to start the cloud data processing flow: First, it determines the total number of transmitted frame blocks required to constitute the junction temperature data packet; the total number of transmitted frame blocks N is determined by whether fragmentation is used and the fragment size. Next, it assigns a weight coefficient to each transmitted frame block, with the following allocation rule: the weight coefficient of a transmitted frame block containing complete data of the digital temperature value field in the payload is set to a first fixed value, and the weight coefficient of a negative value field is set to a first fixed value. For transmission frame blocks that only contain a timestamp field and a cyclic redundancy check (CRC) field, the sum of their weight coefficients is a second fixed value. The weight coefficient of each such transmission frame block is the second fixed value divided by the number of such transmission frame blocks. The first fixed value is set to 0.6, and the second fixed value is set to 0.4, reflecting the core importance of digital temperature data. If multiple non-core frame blocks exist, the weight of each is 0.4 divided by the number of non-core frame blocks. Then, a binary reception status flag is defined for each transmission frame block. If the transmission frame block is successfully received and passes the check, its flag is the first status value; otherwise, it is the second status value. The first status value is set to 1, indicating success; the second status value is set to 0, indicating failure.

[0099] Next, the weighted success rate is calculated by summing the products of the weight coefficients of all transmitted frame blocks and their corresponding reception status flags. Finally, the calculated weighted success rate is compared with a preset success reception threshold, which is a decimal between 0 and 1. The preset success reception threshold is 0.95, meaning that a small amount of non-core data loss is allowed, but core temperature data loss is not permitted, to balance reliability and processing timeliness. Only when the weighted success rate is greater than or equal to this success reception threshold will the cloud server automatically generate a data processing flow trigger with a valid status. Send a signal; if the weighted success rate is less than the success reception threshold, a trigger signal with an invalid value is generated or not generated. The valid trigger signal is a logic high level or a specific event message, and the invalid signal is a logic low level or no signal. The trigger signal will wake up the cloud data analysis service. The data processing flow trigger signal with a valid value is used to start the cloud data processing flow for the timestamp and digital temperature value contained in the junction temperature data packet. The started data processing flow includes: storing the timestamp and digital temperature value into the historical database, calling the machine learning model to perform trend analysis, and generating subsequent control instructions.

[0100] In this embodiment, it is specifically necessary to explain that the process of performing trend analysis based on the received junction temperature data packet and calling the historical temperature database in the cloud analysis module includes two stages: data fusion and cleaning, and spatiotemporal feature extraction. Specifically:

[0101] During the data fusion and cleaning phase, the cloud analysis module receives junction temperature data packets from the data communication module and extracts the current timestamp and digital temperature value. Based on the extracted current timestamp and digital temperature value, it queries the historical temperature database to obtain the historical temperature sequence of the same optical module device within a preset historical time window. This time window is a fixed duration traced back from the current timestamp. The length of the preset historical time window is configurable, with a typical value of 60 seconds. This duration is sufficient to cover the typical thermal transient process caused by load changes in the optical module, providing sufficient contextual information for trend analysis. At the same time, it obtains the set of digital temperature values ​​of other optical module devices deployed in the same equipment rack and with the same topological location near the current timestamp. "Nearby" usually means a time deviation within ±100 milliseconds to ensure that spatial correlation analysis is based on the thermal state at approximately the same moment. The obtained set of neighboring device temperatures is used to construct a spatial thermal coupling map.

[0102] Flow cytometry anomaly cleaning is performed on historical temperature sequences using a modified Grubbs test. For each digital temperature value in the historical temperature sequence designated as a check point, the following calculation and judgment steps are performed sequentially: D1. Calculate the time-weighted moving median centered on the check point. The calculation process is as follows: Assign a weight value to each other digital temperature value within a time window before and after the check point in the historical temperature sequence. This weight value is calculated by dividing the absolute value of the difference between the timestamp of the other digital temperature value and the timestamp of the check point by a preset time decay constant, taking the negative value, and then performing an exponential operation with the natural constant e as the base. Multiply all other digital temperature values ​​within the window by their corresponding weights, sum them, and then divide by the sum of all weight values. The result is the time-weighted moving median. The typical value of the time decay constant is 1.0 second. This weight calculation method assigns higher weights to data points closer to the check point, so that the moving median can... To more sensitively reflect the latest trends and resist sudden jumps; D2, calculate a robust scaling estimate. The calculation process is as follows: first, calculate the absolute difference between each other digital temperature value within the time window and the time-weighted moving median obtained in step one; then, find the median from these absolute differences; finally, multiply the median by a constant 1.4826. The result is the robust scaling estimate. The constant 1.4826 is used to convert the absolute deviation of the median into a consistent estimate of the standard deviation of the normal distribution, thus making the robust scaling estimate statistically significant and insensitive to outliers; D3, calculate the standardized deviation of the point to be inspected. The calculation process is as follows: first, calculate the absolute difference between the digital temperature value of the point to be inspected and the time-weighted moving median obtained in D1; then, divide this absolute difference by the robust scaling estimate obtained in D2. The result is the standardized deviation. The standardized deviation makes the original deviation dimensionless, so that data points from different temperature ranges can be uniformly judged for anomalies;

[0103] After the calculation is completed, a judgment is made: if the calculated standardized deviation value is greater than a critical value, the temperature value of the test point is determined to be an outlier and is removed or replaced by a smoothing function using an adjacent normal value. The critical value is based on a preset significance level and the number of valid samples actually involved in the weighted calculation in the first step. It is obtained by consulting the Grubbs test critical value table. The preset significance level is usually set to 0.05, corresponding to a 95% confidence level. The number of valid samples is the number of data points that exceed the total weight and a certain proportion (such as 90%). The critical value decreases as the number of valid samples increases, ensuring the rigor of the test.

[0104] After cleaning, the cleaned historical temperature sequence, the set of digital temperature values ​​from nearby devices, and the current timestamp and digital temperature value are merged to form a multidimensional input dataset.

[0105] During the spatiotemporal feature extraction stage, the following operations are performed:

[0106] First, time attention feature extraction is performed on the cleaned historical temperature sequence. Specifically, for each numerical temperature value in the historical temperature sequence, the encoded vector obtained by mapping the numerical temperature value and its corresponding timestamp through an embedding function is concatenated to form a combined feature vector. The timestamp embedding function converts the absolute timestamp into periodic sine and cosine codes to capture potential thermal behavior patterns at different times or during different operating cycles within a day. To calculate the attention level of any target point to another historical point in the historical temperature sequence, an attention weight calculation process is defined: the combined feature vector of the target point is multiplied by a learnable first weight matrix to obtain the query vector; the combined feature vector of the historical point is multiplied by a learnable second weight matrix to obtain the key vector; the query vector and the key vector are concatenated sequentially into a long vector, and then this long vector is multiplied by a learnable weight vector. The resulting scalar is nonlinearly transformed through a LeakyReLU activation function. The transformation result is then exponentially calculated with the natural constant e as the base to obtain the unnormalized attention score of the historical point to the target point. The negative slope parameter of the LeakyReLU activation function is typically set to 0.01 to alleviate the gradient vanishing problem. After calculating the unnormalized attention score of all historical points in the historical temperature sequence in the same way, all scores are summed, and then the unnormalized attention score of each historical point is divided by the sum. The quotient is the final attention weight of the historical point to the target point. This final attention weight is a value between zero and one. After calculating the attention weights between all points in the historical temperature sequence, these weights are used to perform a weighted summation of the combined feature vectors of all historical points to generate a time feature vector that reflects the dynamic changes of the target device's own temperature over time. The dimension of the time feature vector generated by the weighted summation can be determined through model training, with typical values ​​of 64 or 128 dimensions, which can condense the key time patterns in the historical temperature sequence.

[0107] Simultaneously, spatial graph convolution feature extraction is performed to process the set of digital temperature values ​​of neighboring devices. Specifically, each neighboring optical module device is treated as a node in the graph, and the feature of this node is its digital temperature value. An edge is established between any two nodes, and the weight of this edge is determined by the physical distance between the two nodes and a predefined heat dissipation coupling coefficient. The weight decreases as the physical distance increases and increases as the heat dissipation coupling coefficient increases. Based on all nodes and edges, a graph structure is constructed, and an adjacency matrix is ​​used to represent this graph structure. The value of each element in the adjacency matrix is ​​the weight of the edge between the corresponding node pair. The sum of this adjacency matrix and an identity matrix is ​​calculated to obtain an augmented adjacency matrix with self-connections. The sum of the elements in each row of the augmented adjacency matrix is ​​calculated, and a diagonal matrix with these sums as its diagonal elements is constructed. The negative 1 / 2 power of the diagonal matrix is ​​multiplied by the augmented adjacency matrix, and then multiplied by another negative 1 / 2 power of the diagonal matrix to obtain a normalized matrix for feature propagation. The digital temperature values ​​of all nodes are arranged into a feature matrix. The normalized matrix is ​​multiplied by the feature matrix, and the result is multiplied by a learnable weight matrix. Finally, the result is passed through a non-linear activation function to generate a new feature vector for each node that aggregates information from its neighboring nodes. The non-linear activation function is usually the ReLU function. The size of the learnable weight matrix determines the dimension of the output feature vector, for example, mapping from 1 dimension (original temperature) to 16 dimensions. From these new feature vectors, a new feature vector representing the target optical module device is extracted as a spatial feature vector reflecting the thermal interference of the target optical module device from surrounding devices.

[0108] Finally, the temporal feature vector and the spatial feature vector are fused to obtain a high-level fused feature tensor containing spatiotemporal context information. The fusion method is either vector concatenation or element-wise addition. If concatenation is used, the dimension of the high-level fused feature tensor is the sum of the dimensions of the temporal feature vector and the spatial feature vector.

[0109] The specific process of generating and outputting control signals using machine learning algorithms includes two steps: probabilistic sequence prediction and robust control optimization. Specifically:

[0110] In the probabilistic sequence prediction step, a variational Bayesian long short-term memory (LSTM) network model is used to process the high-level fusion feature tensor to predict the junction temperature change trajectory and its uncertainty within a preset time window. The preset time window is typically 5 to 30 seconds, depending on the thermal time constant and control response speed of the cooling system. The prediction principle of the variational Bayesian LSTM network model is based on Bayesian inference, specifically: all adjustable weight parameters within the variational Bayesian LSTM network model are considered as a whole, denoted as weight random variables; to describe the uncertainty of this weight random variable, a probability distribution determined by its variational parameters is defined, called the variational posterior distribution of the weight random variable; the training process of the variational Bayesian LSTM network model involves learning and optimizing these variational parameters so that the variational posterior distribution is as close as possible to the true posterior distribution of the weight random variable under given training data; during prediction, for a given high-level fusion feature tensor input... The predicted output of the variational Bayesian long short-term memory network model is no longer a single numerical sequence, but a probability distribution of the future junction temperature sequence. This probability distribution is obtained as follows: considering all possible values ​​of all weighted random variables under their variational posterior distribution, for each possible weight value, the variational Bayesian long short-term memory network model performs a forward propagation calculation with the high-level fusion feature tensor as input, obtaining a deterministic prediction of the future junction temperature sequence; finally, the numerous deterministic prediction results generated by all possible weight values ​​are weighted and averaged or integrated according to their variational posterior distribution, thus obtaining the final output probability distribution of the future junction temperature; and then it outputs two sequences: the mean sequence of the predicted junction temperature at each future time point, and the standard deviation sequence representing the prediction uncertainty. In practice, the weighted average or integration is approximated by multiple Monte Carlo sampling, for example, performing 50 to 100 forward propagation calculations and taking the mean and standard deviation of the results;

[0111] In the robust control optimization step, a stochastic model predictive control problem is constructed, simplifying the optical module heat dissipation system into a first-order thermodynamic model. Using the junction temperature prediction mean sequence and the standard deviation sequence representing prediction uncertainty output from the probabilistic sequence prediction step as inputs, an optimization objective function is defined. The construction and solution process of this objective function is as follows:

[0112] First, define a set of future control action sequences containing K consecutive control actions, where each action represents a control command to be determined, such as the duty cycle value of a pulse width modulation signal or the opening value of a coolant valve;

[0113] Secondly, an optimization objective function is constructed. This objective function finds an optimal sequence from all possible future control action sequences that minimizes the overall cost. The overall cost is defined as the weighted sum of the expected temperature penalty cost and the expected power consumption cost within a prediction window containing K time steps in the future. Its specific construction process is as follows:

[0114] The calculation of the expected temperature penalty cost is based on the probability distribution of future junction temperatures. Specifically, the expected value of the temperature penalty term is calculated for each future time step. The temperature penalty term is constructed as a function of the predicted future junction temperature, and its key characteristic is that when the predicted junction temperature approaches or exceeds a preset junction temperature safety threshold (e.g., 85 degrees Celsius), the value of the penalty term increases sharply in a non-linear manner (e.g., designed as a piecewise function or exponential function), thereby significantly increasing the cost of over-temperature risk in the optimization objective, so as to strictly avoid the situation where the junction temperature exceeds the safe range.

[0115] The calculation of the expected power consumption cost is also based on the mathematical expectation of the power consumption penalty term for each future time step. The power consumption penalty term is a function of the future control action and is typically designed to be proportional to the square of the magnitude of the control action (e.g., proportional to the square of the fan speed or the square of the drive current of the thermoelectric cooler) to encourage a smooth and energy-efficient control strategy.

[0116] Then, the expected temperature penalty cost is multiplied by a positive first weighting coefficient, and the expected power consumption cost is multiplied by a positive second weighting coefficient. These two values ​​are then added together to obtain the weighted cost for that time step. Finally, the weighted costs for all K time steps within the prediction window are summed, and the resulting sum is the value of the optimization objective function that needs to be minimized. The first and second weighting coefficients are trade-off constants used to balance temperature control performance and system energy consumption. Their typical values ​​range from 0.01 to 0.1 and can be tuned using trial and error or more advanced optimization methods (such as reinforcement learning).

[0117] The essence of this construction method is to explicitly incorporate the prediction uncertainty (i.e., the standard deviation sequence) output by the probabilistic sequence prediction step into the evaluation of the control strategy through mathematical expectation calculation. This allows the optimization process to consider not only the most likely future temperature trend, but also all possible trends and their probabilities, thus automatically favoring the selection of a control strategy that remains robust even in the "worst-case" situation, thereby improving the overall reliability of the system.

[0118] Next, the optimization objective is defined: the optimization objective is to find an optimal sequence from all possible future control action sequences that minimizes the value of the optimization objective function;

[0119] Then, the optimization solution is performed: the defined optimization objective function is minimized using a stochastic optimization algorithm to calculate the optimal future control action sequence. The stochastic optimization algorithm can use stochastic gradient descent, covariance matrix adaptive evolution strategy or particle swarm optimization, etc., to iteratively find the control sequence that minimizes the objective function in an online or offline environment.

[0120] Finally, output control commands: extract the first control action from the optimal future control action sequence obtained from the solution, as an immediate control command that should be issued to the local execution module immediately. This immediate control command is quantified into a specific value that the local execution module can recognize. For example, for a fan, it may be the target speed percentage; for a thermoelectric cooler, it may be the drive current setpoint.

[0121] The real-time control command is encapsulated into a data packet of control signals and sent to the local execution module via the communication link.

[0122] In this embodiment, the specific process of receiving control signals from the cloud in the local execution module is as follows:

[0123] The system receives control signals encapsulated in data packets from the cloud analysis module via a communication interface. First, it decrypts and verifies the integrity of the data packet. Decryption uses the AES-128 algorithm, and integrity verification uses the CRC-16 algorithm. The decryption key and verification seed are preset by the system. Upon successful decryption, the real-time control command is extracted. Simultaneously, the local execution module synchronously acquires the latest digital temperature value from the temperature monitoring module. The acquisition period for the digital temperature value is synchronized with the local closed-loop fine-tuning cycle, for example, 1 millisecond, to achieve rapid response. Then, it performs dynamic limiting operations at the safety boundary: based on the proximity of the digital temperature value to a preset junction temperature safety threshold, it dynamically calculates a lower limit and an upper limit for a safe execution command. The junction temperature safety threshold is preset to 85 degrees Celsius. The quantification method for "proximity" is: calculating the absolute difference between the digital temperature value and the junction temperature safety threshold, defining a temperature difference threshold, for example, 5 degrees Celsius. When the absolute difference is less than or equal to this temperature difference threshold... When the value is close to the threshold, it is determined to be "close". At this time, the lower limit of the safety execution command should be dynamically increased. The lower limit and upper limit of the safety execution command constitute a dynamic safety execution range. The rule is: the closer the digital temperature value is to the junction temperature safety threshold, the higher the lower limit of the safety execution command, while the upper limit remains at the maximum allowable value or the value set according to the protection strategy. Then, a limit calculation is performed. This calculation receives three inputs: the parsed immediate control command, the lower limit of the safety execution command, and the upper limit of the safety execution command. The immediate control command is compared with the lower limit and the upper limit of the safety execution command respectively. If it is lower than the lower limit of the safety execution command, the output value is equal to the lower limit of the safety execution command; if it is higher than the upper limit of the safety execution command, the output value is equal to the upper limit of the safety execution command; if it is between the lower limit and the upper limit of the safety execution command, the output value remains unchanged, which is the immediate control command itself. The output result of this limit calculation is the control command to be executed after local safety verification.

[0124] The specific process of changing the heat dissipation state of the optical module by adjusting the current of the thermoelectric cooler or the fan speed through the drive circuit is as follows:

[0125] First, drive quantity mapping is performed: based on the normalized cooling demand intensity represented by the control command to be executed, it is linearly converted into a requested cooling power value in units of power; and combined with the preset inverse characteristic model of the thermoelectric cooler or fan, a preliminary drive setpoint is calculated. The inverse characteristic model is a data lookup table, denoted as the inverse mapping function. Its inputs include the requested cooling power value and the current operating point parameters of the actuator. For thermoelectric coolers, the operating point parameters include the temperatures of their hot and cold ends; for fans, they include the static air pressure at their inlet. The data lookup table of the inverse characteristic model is generated during the factory calibration phase, calibrated in a standard environment (e.g., 25 degrees Celsius room temperature, 50% relative humidity). The process is conducted under humidity conditions, applying a drive signal in fixed steps (e.g., every 0.1 amperes or every 1% duty cycle) to measure steady-state cooling power, thereby establishing a mapping relationship between "drive signal - cooling power - operating point". When in use, this lookup table is used for bilinear interpolation based on the current operating point parameters. The update or query frequency of the lookup table is consistent with the drive cycle, for example, once every 10 milliseconds. The output of the inverse mapping function is the initial drive setting value; for thermoelectric coolers, this value is the drive current setting value, and for fans, it is the pulse width modulation duty cycle setting value. The inverse characteristic model describes the nonlinear mapping relationship between cooling demand intensity and physical drive quantity, and its parameters are obtained through pre-calibration.

[0126] Next, anti-saturation dynamic compensation is performed: To prevent the control performance from deteriorating due to the drive circuit or actuator reaching its physical limits, a model reference adaptive anti-saturation compensator is introduced. This model reference adaptive anti-saturation compensator takes the initial drive setpoint as an ideal input and calculates and outputs a final drive setpoint after anti-saturation compensation in real time through an adaptive adjustment law. The mathematical expression of the adaptive adjustment law describes the rate of change of the final drive setpoint over time, which consists of two terms: the first term is the tracking term, whose value is a constant called the tracking gain multiplied by the difference between the initial drive setpoint and the final drive setpoint, used to drive the final drive setpoint to track the initial drive setpoint; the second term is the anti-saturation compensation term, whose value is a constant called the anti-saturation compensation gain multiplied by the output of a saturation function, the output of which... Input is the difference between the final drive setpoint and a preset physical drive limit. The saturation function outputs the input value when the absolute value of the input is less than a small positive boundary value, and outputs the sign of the boundary value multiplied by the boundary value when the absolute value of the input is greater than or equal to the boundary value. The typical value for tracking gain is 0.5, and the typical value for anti-saturation compensation gain is 0.3. The physical drive limit is determined by the actuator specifications, for example, a TEC with a maximum current of 5 amps and a fan with a maximum PWM duty cycle of 100%. The boundary value of the saturation function is set to ±10% of the drive limit. For example, when the drive value enters the range of 90% to 100% of the maximum value, the anti-saturation compensation term takes effect, making it smoothly approach rather than hard-hitting the limit value. This adjustment law allows the final drive setpoint to decelerate smoothly in advance when it approaches the physical limit, thereby avoiding hard saturation.

[0127] When both thermoelectric coolers and fans are configured as actuators, a multi-actuator coordination strategy is implemented. This strategy includes dynamic weight allocation calculation, and the specific process is as follows:

[0128] T1. Based on the digital temperature value, the energy efficiency ratio coefficient of the thermoelectric cooler and the fan at the current temperature are obtained through a preset lookup table. The energy efficiency ratio coefficient represents the cooling power that can be provided by a unit of electrical power consumption. The energy efficiency ratio coefficient lookup table is established based on the calibration experimental data. The temperature range is usually divided into 10 degrees Celsius levels. The coefficient calculation formula can be obtained by fitting the experimental data based on the thermodynamic model, reflecting the change of the heat dissipation efficiency of the actuator at different temperatures.

[0129] T2. Calculate a dynamic allocation coefficient. The numerator of this coefficient is the energy efficiency ratio (EER) of the thermoelectric cooler, and the denominator is the sum of the EER of the thermoelectric cooler, the EER of the fan, and a positive constant to prevent division by zero. Using this dynamic allocation coefficient, the normalized cooling demand intensity is decomposed into a cooling demand component for the thermoelectric cooler and a cooling demand component for the fan. Specifically, the cooling demand component for the thermoelectric cooler is equal to the requested cooling power value multiplied by the dynamic allocation coefficient, and the cooling demand component for the fan is equal to the requested cooling power value multiplied by one minus the difference of the dynamic allocation coefficient. The final drive setpoints are calculated separately using their respective inverse characteristic models and anti-saturation compensators. In multi-actuator collaboration, a priority strategy can be set. For example, when the digital temperature value is extremely high, priority is given to ensuring that the thermoelectric cooler reaches its maximum capacity, and then the remaining demand is allocated to the fan.

[0130] Finally, local closed-loop fine-tuning and drive are performed: the final drive setpoint is converted into a high-precision analog current signal or pulse width modulation signal and applied to the drive circuit of the thermoelectric cooler or fan; simultaneously, local fast closed-loop fine-tuning is performed to suppress high-frequency thermal disturbances. This local fast closed-loop fine-tuning process is based on the calculation and superposition of a fine-tuning amount, specifically: first, a temperature deviation value is calculated, which is the difference between a desired temperature derived from the cooling requirements implied in the control command to be executed and a digital temperature value. The desired temperature can be derived from the control command to be executed through a preset simple linear or lookup table model that reflects the relationship between cooling intensity and target temperature. For example, a control command of 0.5 may correspond to a desired temperature of 70 degrees Celsius; simultaneously, the rate of change of the temperature deviation value over time is calculated; then, based on a high-speed, low-gain proportional-derivative control law combined with a nonlinear compensation term, the fine-tuning amount is calculated. The specific composition of the control law and the nonlinear compensation term is: the fine-tuning amount equals a first proportional constant multiplied by the temperature deviation value, plus a first differential constant multiplied by the temperature deviation value. The rate of change is then subtracted from a nonlinear compensation term. The nonlinear compensation term equals a second compensation constant multiplied by the sign function of the temperature deviation value, and then multiplied by the power of the absolute value of the temperature deviation value. The exponent of the power is a third constant between zero and one. The first proportional constant can be set to 0.1, the first differential constant can be set to 0.05, the second compensation constant can be set to 0.2, and the third constant (i.e., the exponent m) is fixed at 0.5 to achieve superlinear response characteristics. The sign function sign(e) is defined as follows: when the temperature deviation value e is greater than zero, the function value is 1; when e is equal to zero, the function value is 0; when e is less than zero, the function value is -1. The first proportional constant and the first differential constant are both set to small positive values ​​to suppress high-frequency small-amplitude disturbances. The second compensation constant and the third constant between zero and one are used together to provide additional fast compensation force with superlinear response characteristics when large and rapid temperature changes occur. Finally, the calculated fine-tuning amount is dynamically superimposed on the final drive setpoint to form a composite control signal acting on the drive circuit.

[0131] After completing the drive and fine-tuning, the execution status information is encapsulated: the currently applied final drive setpoint, fine-tuning amount, real-time acquired digital temperature value, and local controller's operating status flags are encapsulated into a status feedback data packet. The operating status flags may include binary flags such as "anti-saturation compensation activated," "safety limit effective," and "fine-tuning amount exceeded." The status feedback data packet contains a fixed-format header, various data fields, and a checksum. This status feedback data packet is sent to the cloud analysis module via the communication link for online learning and adaptive optimization of the cloud model. The status feedback data packet is reported using the uplink idle time slot of the data communication module or the next data transmission cycle. The cloud analysis module uses this real execution and status data to perform online correction and optimization of its internal thermodynamic model parameters, prediction model, and control strategy.

[0132] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0133] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0134] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0135] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The function specified in one or more boxes.

[0136] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0137] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0138] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A cloud-based collaborative control system for the junction temperature of an optical module, characterized in that, Specifically, it includes: The temperature monitoring module, data communication module, cloud analysis module, and local execution module are connected in sequence, among which; Temperature monitoring module: When the optical module is powered on, it collects the junction temperature analog signal through the temperature sensor integrated near the optical module chip, converts the junction temperature analog signal into a digital temperature value, and generates a junction temperature data packet; Data communication module: Receives junction temperature data packets, encapsulates the junction temperature data packets into data frames through the wireless network protocol and uploads them to the cloud server, triggering the cloud data processing flow; Cloud analytics module: Deployed on a cloud server, it calls the historical temperature database based on the received junction temperature data packets to perform trend analysis, and uses machine learning algorithms to generate and output control signals; The specific process of performing trend analysis based on the received junction temperature data packet and historical temperature database includes two stages: data fusion and cleaning, and spatiotemporal feature extraction. During the data fusion and cleaning phase, the timestamp and digital temperature value of the current moment are extracted, the historical temperature database is queried, and the historical temperature sequence of the same optical module device within a preset historical time window is obtained; at the same time, the set of digital temperature values ​​of other optical module devices deployed in the same equipment rack and with the same topological location near the current timestamp is obtained. Flow cytometry anomaly cleaning is performed on historical temperature sequences. This process uses a modified Grubbs test method. For each digital temperature value in the historical temperature sequence that serves as a check point, the following calculation and judgment steps are performed sequentially: D1. Calculate the time-weighted moving median centered on the check point. The calculation process is as follows: Assign a weight value to each other digital temperature value in the historical temperature sequence located within a time window before and after the check point. This weight value is obtained by dividing the absolute value of the difference between the timestamp of the other digital temperature value and the timestamp of the check point by a preset time decay constant, taking the negative number, and then performing an exponential operation with the natural constant e as the base; Multiply each of the other digital temperature values ​​located within the window by its corresponding weight, sum them up, and then divide by the sum of all weight values. The result is the time-weighted moving median. D2. Calculate a robust scaling estimate. The calculation process is as follows: First, calculate the absolute difference between each other numerical temperature value within the time window and the time-weighted moving median obtained in the first step. Then, find the median from these absolute differences. Finally, multiply the median by a constant of 1.4826. The result is the robust scaling estimate. D3. Calculate the standardized deviation of the test point. The calculation process is as follows: First, calculate the absolute difference between the digital temperature value of the test point and the time-weighted moving median obtained in D1. Then, divide this absolute difference by the robust scaling estimate obtained in D2. The result is the standardized deviation. After the calculation is completed, a judgment is made: if the calculated standardized deviation value is greater than a critical value, the digital temperature value of the point to be inspected is determined to be an outlier and is removed or replaced by a smoothing function with an adjacent normal value. After cleaning, the cleaned historical temperature sequence, the set of digital temperature values ​​from nearby devices, and the current timestamp and digital temperature value are merged to form a multidimensional input dataset. In the spatiotemporal feature extraction stage, temporal attention feature extraction is performed to process the cleaned historical temperature sequence and generate a temporal feature vector; at the same time, spatial graph convolution feature extraction is performed to process the set of digital temperature values ​​of nearby devices and generate a spatial feature vector; the temporal feature vector and the spatial feature vector are fused into a high-level fusion feature tensor. The specific process of generating and outputting control signals using machine learning algorithms includes two steps: probabilistic sequence prediction and robust control optimization. Specifically: In the probabilistic sequence prediction step, a variational Bayesian long short-term memory network model is used to process the advanced fusion feature tensor, and output the predicted mean and standard deviation sequences of the junction temperature at each future time point. In the robust control optimization step, a stochastic model predictive control problem is constructed, simplifying the optical module heat dissipation system into a first-order thermodynamic model. Using the junction temperature prediction mean sequence and standard deviation sequence as input, an optimization objective function is defined as the weighted sum of the expected temperature penalty cost and the expected power consumption cost within the future prediction window. Then, optimization is performed: a stochastic optimization algorithm is used to minimize the defined optimization objective function, calculating the optimal future control action sequence. Finally, output control instructions: extract the first control action from the optimal future control action sequence obtained from the solution, and use it as an immediate control instruction that should be issued to the local execution module immediately; The real-time control command is encapsulated into a data packet of control signals and sent to the local execution module via the communication link; Local execution module: Receives control signals from the cloud, adjusts the current of the thermoelectric cooler or the fan speed through the drive circuit, changes the heat dissipation state of the optical module, and stabilizes the junction temperature within the preset threshold range.

2. The cloud-based collaborative control system for optical module junction temperature according to claim 1, characterized in that: In the temperature monitoring module, the specific operation of acquiring the junction temperature analog signal through the temperature sensor integrated near the optical module chip is as follows: The temperature sensor uses a surface-mount negative temperature coefficient thermistor, which is tightly mounted on the surface of the optical module chip package. When the optical module is powered on, the resistance value of the thermistor changes with the junction temperature. A constant current is provided by the constant current source circuit to generate an initial junction temperature analog signal. Within the sampling time window, multiple instantaneous voltage sample values ​​are continuously acquired. Each sample value is assigned an exponentially decaying weight coefficient, and a weighted average analog voltage signal is calculated. The temperature monitoring module also includes a signal conditioning circuit, which contains an operational amplifier to amplify the weighted average analog voltage signal to obtain an analog voltage signal as the input for analog-to-digital conversion.

3. The cloud-based collaborative control system for optical module junction temperature according to claim 2, characterized in that: The specific process of converting the junction temperature analog signal into a digital temperature value to generate a junction temperature data packet is as follows: The analog voltage signal, which serves as the input for analog-to-digital conversion, is input to the analog-to-digital converter (ADC) to obtain the original digital code. A ternary polynomial calculation process is used to map the original digital code into a digital temperature value. The digital temperature value is stored in the form of a 16-bit unsigned integer. The conversion result is then digitally filtered to obtain the filtered digital temperature value. Finally, it is used to assemble the junction temperature data packet, which includes a data packet header, a timestamp field, a digital temperature value field, and a cyclic redundancy check (CRC) field.

4. The cloud-based collaborative control system for optical module junction temperature according to claim 3, characterized in that: In the data communication module, the specific operation of encapsulating junction temperature data packets into data frames using a wireless network protocol is as follows: The process of receiving junction temperature data packets and encapsulating them into data frames involves sequentially performing frame header construction, payload mapping, and frame assembly. The data frame control header is generated according to the link layer protocol. The data frame control header contains a data feature identifier field. In addition to the source address, destination address and protocol type fields, the frame control header also adds a data feature identifier field. The data feature identifier is calculated based on the timestamp field and the digital temperature value field using a lightweight hash function; the junction temperature data packet is used as the payload of the data frame and combined with the frame control header to form the primary data frame.

5. The cloud-based collaborative control system for optical module junction temperature according to claim 4, characterized in that: The specific process for triggering the cloud data processing flow is as follows: Before sending the primary data frame, the communication quality of the current wireless channel is monitored in real time to obtain a signal-to-noise ratio (SNR) estimate that characterizes the channel quality. Based on the comparison between the SNR estimate and a preset threshold, it is decided whether to fragment the primary data frame into transmission frame blocks. The dynamic coding rate of the forward error correction coding is dynamically calculated. The transmission frame blocks are then channel-coded using the dynamic coding rate and sent. The cloud server decodes and reassembles the received transmission frame blocks to recover the junction temperature data packets; Calculate the weighted success rate of all transmitted frame blocks; trigger the cloud data processing flow only when the weighted success rate is greater than or equal to the preset success rate threshold.

6. The cloud-based collaborative control system for optical module junction temperature according to claim 5, characterized in that: The specific process of receiving control signals from the cloud in the local execution module is as follows: The system receives data packets of control signals and parses out real-time control commands; the synchronous local execution module synchronously obtains the latest digital temperature value from the temperature monitoring module; and dynamically calculates the lower and upper limits of the safe execution command based on the proximity of the digital temperature value to the junction temperature safety threshold. By limiting the amplitude calculation, the real-time control commands are restricted to a safe execution range, and the control commands to be executed are output.

7. The cloud-based collaborative control system for optical module junction temperature according to claim 6, characterized in that: The specific process of changing the heat dissipation state of the optical module by adjusting the current of the thermoelectric cooler or the fan speed through the drive circuit is as follows: First, perform drive quantity mapping: based on the normalized cooling demand intensity represented by the control command to be executed, linearly convert it into a requested cooling power value in units of power; and combine it with the preset inverse characteristic model of the thermoelectric cooler or fan to calculate a preliminary drive setting value. Next, anti-saturation dynamic compensation is performed: a model reference adaptive anti-saturation compensator is introduced; this model reference adaptive anti-saturation compensator takes the initial drive setting value as the ideal input, and calculates and outputs a final drive setting value after anti-saturation compensation in real time through an adaptive adjustment law. When both thermoelectric coolers and fans are configured as actuators, a multi-actuator coordination strategy is implemented. This strategy includes dynamic weight allocation calculation, and the specific process is as follows: T1. Based on the digital temperature value, the energy efficiency ratio coefficient of the thermoelectric cooler and the energy efficiency ratio coefficient of the fan at the current temperature are obtained through a preset lookup table. T2. Calculate a dynamic allocation coefficient. The numerator of this dynamic allocation coefficient is the energy efficiency ratio coefficient of the thermoelectric cooler, and the denominator is the sum of the energy efficiency ratio coefficient of the thermoelectric cooler, the energy efficiency ratio coefficient of the fan, and a positive constant. Using this dynamic allocation coefficient, the normalized cooling demand intensity is decomposed into a cooling demand component for the thermoelectric cooler and a cooling demand component for the fan. Specifically, the cooling demand component for the thermoelectric cooler is equal to the requested cooling power value multiplied by the dynamic allocation coefficient, and the cooling demand component for the fan is equal to the requested cooling power value multiplied by one minus the difference of the dynamic allocation coefficient. And calculate their respective final drive setpoints through their respective inverse characteristic models and anti-saturation compensators. Finally, local closed-loop fine-tuning and drive are performed: the final drive setting value is converted into a high-precision analog current signal or pulse width modulation signal and applied to the drive circuit of the thermoelectric cooler or fan; at the same time, local fast closed-loop fine-tuning is performed. This local fast closed-loop fine-tuning process is based on the calculation and superposition of a fine-tuning amount, specifically: first, a temperature deviation value is calculated, which is the difference between a desired temperature derived from the cooling requirements implied by the control command to be executed and a digital temperature value; at the same time, the rate of change of this temperature deviation value over time is calculated. Then, based on a high-speed, low-gain proportional-derivative control law and combined with a nonlinear compensation term, the fine-tuning amount is calculated; finally, the calculated fine-tuning amount is dynamically superimposed on the final drive setpoint to form a composite control signal acting on the drive circuit. After completing the drive and fine-tuning, the status information is encapsulated: the final drive setting value actually applied, the fine-tuning amount, the real-time acquired digital temperature value, and the local controller's operating status flag are encapsulated together into a status feedback data packet. The status feedback data packet is sent to the cloud analysis module via the communication link for online learning and adaptive optimization of the cloud model.

8. A cloud-based collaborative control method for optical module junction temperature, applied to a cloud-based collaborative control system for optical module junction temperature as described in any one of claims 1-7, characterized in that: Specifically, the following steps are included: Step S1: When the optical module is powered on, a surface-mount negative temperature coefficient thermistor mounted close to the surface of the optical module chip package is used to collect the junction temperature analog signal. The resistance change of the thermistor is converted into a voltage signal through a constant current source circuit to generate an initial junction temperature analog voltage. The initial junction temperature analog voltage is amplified and filtered by a signal conditioning circuit, and then converted into a digital temperature value by an analog-to-digital converter. The digital temperature value is then packaged into a junction temperature data packet containing a data packet header, a timestamp field, a digital temperature value field, and a cyclic redundancy check code field. Step S2: Receive junction temperature data packets through the wireless network communication interface, construct a data frame control header according to the current link layer protocol specification, encapsulate the junction temperature data packets as the payload into a complete data frame, and upload it to the cloud server through the wireless network. Step S3: In the cloud server, parse the junction temperature data packet in the received data frame, call the time series data in the historical temperature database, use machine learning algorithm to analyze the junction temperature change trend, and generate and output control signal for adjusting heat dissipation intensity. Step S4: Locally, the optical module receives control signals from the cloud and converts these signals into corresponding current or pulse width modulation signals through the drive circuit. This adjusts the drive current of the thermoelectric cooler or the speed of the cooling fan, thereby changing the heat dissipation state of the optical module and maintaining the junction temperature within a preset safety threshold range.