Method for determining transmit power and method for training a model for determining transmit power

By using a transmit power determination model and analyzing historical data of multimode optical modules using temporal and spatial neural networks, the problem of transmit power monitoring error in DDM technology is solved, enabling accurate prediction of optical module transmit power and improving the accuracy and reliability of data center network monitoring.

CN121567208BActive Publication Date: 2026-05-05INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INSPUR SUZHOU INTELLIGENT TECH CO LTD
Filing Date
2026-01-22
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing DDM technology has significant errors when monitoring the transmit power of optical modules, resulting in low reliability of monitoring results and failing to meet the high-precision requirements of data center networks.

Method used

By employing a transmit power determination model that combines temporal and spatial neural networks, and analyzing historical operating data of multimode optical modules, multi-channel fusion features are extracted to accurately predict the true transmit power of the optical modules, overcoming errors introduced by temperature drift, component aging, and inter-channel crosstalk.

Benefits of technology

It improves the accuracy and reliability of optical module transmit power prediction, reduces errors, enhances the precision and intelligence of data center network link quality monitoring, and reduces false alarms and missed alarms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121567208B_ABST
    Figure CN121567208B_ABST
Patent Text Reader

Abstract

The application provides a transmission power determination method and a training method of a transmission power determination model, which can be applied to the technical fields of optical communication and artificial intelligence data center. The transmission power determination method comprises the following steps: obtaining historical running data sets of multiple channels of a multi-mode optical module respectively; determining candidate transmission power features of the multiple channels respectively at a current time according to historical running data sequences obtained based on the historical running data sets of the multiple channels respectively by using a time neural network branch of the transmission power determination model; performing feature extraction on a multi-channel feature tensor determined by the multiple historical running data sets by using a spatial neural network branch of the transmission power determination model to obtain multi-channel fusion features; and determining real transmission powers of the multiple channels respectively at the current time according to the multi-channel fusion features and the candidate transmission power features of the multiple channels respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of optical communication technology and artificial intelligence data center technology, and more specifically, to a method for determining transmission power and a training method for a transmission power determination model. Background Technology

[0002] With the rapid development of cloud computing, big data, and artificial intelligence technologies, data center networks are constantly evolving towards higher bandwidth, lower latency, and higher reliability. In data center network systems, optical modules, as core components for optical communication, directly determine the transmission quality of the entire data link based on their operational stability. For example, Digital Diagnostic Monitoring (DDM) technology can be used to monitor the operating parameters of optical modules (such as transmit optical power and receive optical power) to enable quality monitoring and fault warning for these modules.

[0003] However, there is a significant error between the optical module's transmit power monitored using DDM technology and the actual transmit power of the optical module, resulting in low reliability of the transmit power in the monitoring results. Summary of the Invention

[0004] In view of the above problems, this application provides a method for determining transmission power and a method for training a transmission power determination model. Furthermore, this application also provides a transmission power determination apparatus, a training apparatus for a transmission power determination model, equipment, a medium, and a program product.

[0005] According to one aspect of this application, a method for determining the transmit power of a multimode optical module is provided, comprising: acquiring historical operating datasets for multiple channels of the multimode optical module; the historical operating datasets including multiple historical operating data with historical timestamps; using a temporal neural network branch of the transmit power determination model, determining candidate transmit power features for each of the multiple channels at the current moment based on the historical operating data sequence obtained from the historical operating datasets for each of the multiple channels; using a spatial neural network branch of the transmit power determination model, extracting features from the multi-channel feature tensors determined by the multiple historical operating datasets to obtain multi-channel fusion features; the multi-channel fusion features characterizing the single-channel operating state of each of the multiple channels and the inter-channel interference environment between the multiple channels; and determining the true transmit power of each of the multiple channels at the current moment based on the multi-channel fusion features and the candidate transmit power features for each of the multiple channels.

[0006] According to another aspect of this application, a training method for a transmit power determination model for a multimode optical module is provided, comprising: acquiring sample running datasets for multiple sample channels of a sample multimode optical module; the sample running datasets include multiple sample running data with sample timestamps; using a temporal neural network branch of a candidate transmit power determination model, determining the first transmit power feature of each of the multiple sample channels at the current time based on the sample running data sequence obtained from the sample running datasets for each of the multiple sample channels; and using a spatial neural network branch of the candidate transmit power determination model, extracting features from the sample multi-channel feature tensor determined by the multiple sample running datasets to obtain the sample multi-channel... Fusion characteristics; the multi-channel fusion characteristics of the samples represent the single-channel operating status of each of the multiple sample channels and the inter-channel interference environment between the multiple sample channel environments; based on the multi-channel fusion characteristics of the samples and the first transmit power characteristics of each of the multiple sample channels, the predicted transmit power values ​​of each of the multiple sample channels at the current time of the sample are determined; based on the predicted transmit power values ​​of each of the multiple sample channels and the true transmit power values ​​of each of the multiple sample channels, the candidate transmit power determination model is trained to obtain the transmit power determination model; wherein, the true transmit power values ​​of each of the multiple sample channels are: the optical power received by the peer optical module that transmits data with the sample multimode optical module through the multiple sample channels at the current time of the sample.

[0007] Another aspect of this application provides a transmit power determination device for a multimode optical module, comprising: a first acquisition module for acquiring historical operating datasets for multiple channels of the multimode optical module; the historical operating datasets include multiple historical operating data with historical timestamps; a first processing module for using a temporal neural network branch of a transmit power determination model to determine candidate transmit power features for each of the multiple channels at the current moment based on the historical operating data sequence obtained from the historical operating datasets for each of the multiple channels; a second processing module for using a spatial neural network branch of the transmit power determination model to extract features from the multi-channel feature tensors determined by the multiple historical operating datasets to obtain multi-channel fusion features; the multi-channel fusion features characterize the single-channel operating state of each of the multiple channels and the inter-channel interference environment between the multiple channels; and a first determination module for determining the true transmit power of each of the multiple channels at the current moment based on the multi-channel fusion features and the candidate transmit power features for each of the multiple channels.

[0008] Another aspect of this application provides a training apparatus for a transmit power determination model for a multimode optical module, comprising: a second acquisition module for acquiring sample running datasets for multiple sample channels of a sample multimode optical module; the sample running datasets include multiple sample running data with sample timestamps; a third processing module for determining, using a temporal neural network branch of a candidate transmit power determination model, a first transmit power feature of each of the multiple sample channels at the current time based on the sample running data sequence obtained from the sample running datasets for each of the multiple sample channels; and a fourth processing module for extracting features from the multi-channel feature tensor of the sample determined by the multiple sample running datasets using a spatial neural network branch of the candidate transmit power determination model, thereby obtaining the first transmit power feature of each of the multiple sample channels at the current time. The system comprises: a channel fusion feature; a sample multi-channel fusion feature characterizing the single-channel operating state of multiple sample channels and the inter-channel interference environment between multiple sample channel environments; a second determination module, which determines the predicted transmission power of each of the multiple sample channels at the current time based on the sample multi-channel fusion feature and the first transmission power feature of each of the multiple sample channels; and a training module, which trains the candidate transmission power determination model based on the predicted transmission power of each of the multiple sample channels and the actual transmission power of each of the multiple sample channels to obtain the transmission power determination model; wherein, the actual transmission power of each of the multiple sample channels is: the optical power received by the peer optical module that transmits data with the sample multimode optical module through multiple sample channels at the current time of the sample.

[0009] Another aspect of this application provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.

[0010] Another aspect of this application provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.

[0011] Another aspect of this application provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method. Attached Figure Description

[0012] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments of this application with reference to the accompanying drawings.

[0013] Figure 1A The optical power output fitting function of the optical module in the relevant embodiment is shown.

[0014] Figure 1BThe illustration shows an application scenario of the transmit power determination method, apparatus, device, medium, and program product for a multimode optical module according to embodiments of this application.

[0015] Figure 2 A flowchart of a method for determining the transmit power of a multimode optical module according to an embodiment of this application is shown.

[0016] Figure 3A A flowchart illustrating the construction of a multichannel feature tensor according to an embodiment of this application is shown.

[0017] Figure 3B A schematic diagram of constructing a multichannel feature tensor according to an embodiment of this application is shown.

[0018] Figure 4 A flowchart illustrating the determination of candidate transmit power characteristics for each of the multiple channels according to an embodiment of this application is shown.

[0019] Figure 5 A schematic diagram of a fully connected neural network according to an embodiment of this application is shown.

[0020] Figure 6 A schematic diagram of a method for determining the transmit power of a multimode optical module according to an embodiment of this application is shown.

[0021] Figure 7 A flowchart is shown of a training method for a transmit power determination model for a multimode optical module according to an embodiment of this application.

[0022] Figure 8 A structural block diagram of a transmit power determination device for a multimode optical module according to an embodiment of this application is shown.

[0023] Figure 9 A structural block diagram of a training apparatus for determining the transmit power of a multimode optical module according to an embodiment of this application is shown.

[0024] Figure 10 A block diagram of an electronic device suitable for implementing a method for determining the transmit power of a multimode optical module and a method for training a model for determining the transmit power of a multimode optical module, according to an embodiment of this application, is shown. Detailed Implementation

[0025] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.

[0026] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0027] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0028] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0029] In the field of AI training and inference clusters, as the parameters of artificial intelligence models continue to grow, higher demands are placed on computing power and GPU memory. Artificial Intelligence Data Centers (AIDCs) can provide parallel computing based on clusters of multiple GPUs (Graphics Processing Units) to handle the massive computations required for AI model training. However, the GPUs in a data center need to exchange massive amounts of data at high speed, and the actual effective computing power of the cluster depends on the communication efficiency between GPUs, making high-speed interconnection between GPU servers a bottleneck. To address this, optical communication-based interconnection solutions can effectively reduce latency and improve the communication efficiency of GPU clusters, providing solid support for AI computing power.

[0030] Optical communication is a communication method that uses optical signals to transmit information. It uses light as a carrier, modulates various properties of light to load information, and transmits these information-carrying optical signals in a transmission medium. At the receiving end, the optical signals are converted into electrical signals or other forms of signals, thereby realizing information transmission. For example, in data center scenarios, optical communication can be used for high-speed data transmission between servers within a data center, between racks, and between data centers.

[0031] According to embodiments of this application, a data center network may include the aforementioned AIDC. AIDC refers to a comprehensive information processing center that integrates high-performance computing capabilities, artificial intelligence algorithms, and cloud computing services. AIDC can provide computing power, storage, and related services for artificial intelligence and big data applications, such as providing AI model training, inference, data storage, and processing services by adding intelligent computing resources to the data center.

[0032] An optical transceiver module (OTM) is a core component of an optical communication system, responsible for converting between optical and electrical signals. It is widely used in data centers, telecommunications networks, 5G (fifth-generation mobile communication technology), and enterprise networks. By integrating multiple functional components, OTM achieves high-speed data transmission, addressing the needs for high-bandwidth, long-distance, and low-loss interconnection. It can be applied to ultra-large-scale clusters (such as million-chip-level AI data centers).

[0033] In optical modules, a lane refers to an independent channel used for data transmission. A multimode optical module can include multiple lanes, each capable of transmitting data at a certain rate. Multiple lanes can operate in parallel, thereby improving the overall data transmission capacity. For example, a 400 Gbps (gigabits per second) optical module typically uses a multiplexing mode with four lanes.

[0034] In relevant embodiments, DDM technology can be used to monitor the operating parameters of the optical module, such as transmit optical power, receive optical power, optical module operating temperature, laser bias current, and optical module drive voltage. The core of the DDM function is implemented through a microcontroller integrated within the optical module. These microcontrollers can read and process data from various sensors within the optical module and transmit it to the host or network management system via the optical module's electrical interface for processing, enabling functions such as health status detection and fault warning.

[0035] The transmitted optical power monitored using DDM technology is not a true value directly measured by physical sensors, but rather a fitted or estimated value calculated by the microcontroller inside the optical module based on its built-in fitting function (usually calibrated by the optical module manufacturer before shipment). In relevant embodiments, the optical module manufacturer will build an optical power output fitting function (which can be understood as the optical module's transmitted power fitting function) inside the optical module. The optical power can be fitted using the optical module's built-in fitted optical power output function and then displayed through an EEPROM (Electrically Erasable Programmable Read-Only Memory) platform.

[0036] It should be noted that although optical power meters, as instruments specifically designed to measure optical power, can accurately measure the optical power of optical modules, in actual industrial scenarios (especially large-scale online scenarios such as AI clusters), it is impractical to use optical power meters to measure the optical power of a large number of optical modules over a period of time. This would also affect the normal operation of the system due to service interruptions caused by fiber disconnection. Therefore, the industry generally adopts DDM technology to monitor the operating parameters of optical modules.

[0037] Figure 1A The optical power output fitting function of the optical module in the relevant embodiment is shown.

[0038] Figure 1A In the middle, the horizontal axis The vertical axis represents the modulation current of the laser, and the vertical axis represents the optical output power of the optical module. Curves T1 and T2 represent the fitted curves of optical output power at temperature t1 and temperature t2 (t1 < t2), respectively. and These represent the bias currents of the lasers corresponding to curves T1 and T2, respectively. and These represent the modulation current ranges corresponding to curves T1 and T2, respectively. For example, when the temperature is t1, the bias current is... As the modulation current changes, the transmit power of the optical module changes from P0 to P1, and the average output power of the optical module can be expressed as... .like Figure 1A As shown, different temperatures lead to different fitting curves, and the same bias current and modulation current will result in different emitted optical power outputs under different temperatures.

[0039] However, the temperature function curves preset by optical module manufacturers are usually discrete data tables, whose accuracy cannot meet the temperature monitoring requirements of actual cluster deployment environments. The fine-tuning parameters of this output fitting function need to be precisely calibrated with the nominal values ​​of the electronic components inside the optical module. However, in actual industrial scenarios, as the optical module operates for longer periods, the parameter values ​​of its electronic components deviate significantly from the preset values, leading to systematic errors in the optical module transmit power calculated by related methods in large-scale industrial clusters. Specifically, the above-mentioned technical solutions have at least the following inherent errors, causing distortion in the transmit power fitted using the fitting function:

[0040] 1. Temperature Drift Error: The operating temperature of the optical module fluctuates within a wide range, while the parameters of the aforementioned fitting function are usually calibrated at a specific temperature. When the actual operating temperature of the optical module changes, the characteristics of its laser will change, causing the original fitting function to become inaccurate. This results in a significant deviation between the estimated value of the fitted emission power and the actual value of the emission power.

[0041] 2. Component aging and drift error: The electronic components and lasers inside the optical module will age over time, and their performance parameters will gradually drift. The fitting function calibrated at the factory cannot adapt to this slow, long-term change, resulting in the estimation error increasing over time, and the magnitude of the error is difficult to predict.

[0042] 3. Crosstalk between channels exists in multimode optical modules: For multimode optical modules, since the power between lanes will affect each other, the errors caused by temperature drift and aging drift mentioned above will be further amplified compared to single-mode optical modules.

[0043] Therefore, there is an urgent need for a new technology that can overcome the shortcomings of existing DDM technology, which relies on a fixed fitting function, to more accurately and reliably predict or evaluate the true transmission power of optical modules, thereby improving the accuracy and intelligence of data center network link quality monitoring.

[0044] In view of this, embodiments of this application provide a method for determining the transmit power of a multimode optical module. The method includes: acquiring historical operating datasets for multiple channels of the multimode optical module; the historical operating datasets include multiple historical operating data sets with historical timestamps; using a temporal neural network branch of the transmit power determination model, determining candidate transmit power features for each of the multiple channels at the current moment based on the historical operating data sequences obtained from the historical operating datasets for each of the multiple channels; using a spatial neural network branch of the transmit power determination model, extracting features from the multi-channel feature tensors determined by the multiple historical operating datasets to obtain multi-channel fusion features; the multi-channel fusion features characterize the single-channel operating state of each of the multiple channels and the inter-channel interference environment between the multiple channels; and determining the actual transmit power of each of the multiple channels at the current moment based on the multi-channel fusion features and the candidate transmit power features for each of the multiple channels.

[0045] Figure 1B The illustration schematically depicts an application scenario of the transmit power determination method, apparatus, device, medium, and program product for a multimode optical module according to embodiments of this application.

[0046] like Figure 1B As shown, application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0047] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0048] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0049] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0050] It should be noted that the method for determining the transmit power of a multimode optical module provided in this application embodiment can generally be executed by server 105. Correspondingly, the device for determining the transmit power of a multimode optical module provided in this application embodiment can generally be located in server 105. The method for determining the transmit power of a multimode optical module provided in this application embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the device for determining the transmit power of a multimode optical module provided in this application embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.

[0051] It should be understood that Figure 1B The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0052] Figure 2 A flowchart illustrating a method for determining the transmit power of a multimode optical module according to an embodiment of this application is shown.

[0053] like Figure 2 As shown, the method for determining the transmit power of a multimode optical module includes operations S210 to S240.

[0054] In operation S210, the historical operating datasets of each of the multiple channels of the multimode optical module are obtained.

[0055] The historical operation dataset includes multiple historical operation data sets with historical timestamps.

[0056] In operation S220, the time neural network branch of the transmit power determination model is used to determine the candidate transmit power characteristics of each of the multiple channels at the current moment based on the historical operation data sequence obtained from the historical operation datasets of each of the multiple channels.

[0057] In operating S230, the spatial neural network branch of the transmission power determination model is used to extract features from the multi-channel feature tensor determined by multiple historical running datasets, thereby obtaining multi-channel fused features.

[0058] Among them, the multi-channel fusion feature characterizes the single-channel operation status of multiple channels and the inter-channel interference environment between multiple channels.

[0059] In operation S240, the actual transmit power of each of the multiple channels at the current moment is determined based on the multi-channel fusion characteristics and the candidate transmit power characteristics of each channel.

[0060] Multimode optical modules are used for optical communication in data center networks. A multimode optical module includes multiple parallel channels (lanes) for data transmission.

[0061] In one embodiment, historical operational datasets for multiple lanes of a multimode optical module can be collected from a historical operational status database of optical modules. For any given lane, the historical operational dataset includes multiple historical operational data points with historical timestamps. The historical timestamps correspond to historical moments, and the historical operational data can be understood as operational status data at those historical moments.

[0062] As an example, a multimode optical module may include four lanes: lane1, lane2, lane3, and lane4. Historical operational datasets for each of lanes 1, 2, 3, and 4 can be obtained from the optical module's historical operational status database; for example, these could be designated as the first, second, third, and fourth historical operational datasets, respectively. For instance, the first historical operational dataset might include multiple historical operational data points corresponding to lane1, each with its own historical timestamp, and so on.

[0063] According to embodiments of this application, the transmit power determination model can determine the transmit power of the multimode optical module at the current moment based on the historical operating datasets of each of the multiple lanes of the multimode optical module. For example, the transmit power determination model can be a deep learning model trained on a sample dataset related to the multimode optical module.

[0064] Since the transmit power of different lanes in a multimode optical module can influence each other, it is necessary to consider the common characteristics of multiple lanes when predicting the transmit power of the multimode optical module. To address this, the transmit power determination model provided in this application includes a temporal neural network branch and a spatial neural network branch. The temporal neural network branch can be used to analyze the change pattern and dynamic trend of the transmit power of each lane over time, outputting a preliminary predicted value of the transmit power of each lane at the current moment (i.e., candidate transmit power features). The spatial neural network branch can be used to further extract high-dimensional features (i.e., multi-channel fusion features) from the common features of multiple lanes to deeply mine and capture the complex spatial correlations and mutual influences between different channels (lanes). The preliminary predicted value output by the temporal neural network branch can be corrected and fine-tuned based on the multi-channel fusion features extracted by the spatial neural network branch to remove the mutual influence of power between multiple lanes, thereby obtaining the corrected actual predicted value of the transmit power of each lane at the current moment.

[0065] In one embodiment, historical operating data sequences for each of the multiple lanes can be obtained based on their respective historical operating datasets. The temporal neural network branch of the transmit power determination model can be used to determine the candidate transmit power features for each of the multiple lanes at the current moment, based on their respective historical operating data sequences. For any given lane, the candidate transmit power features represent the preliminary prediction of the transmit power of that lane by the temporal neural network branch at the current time.

[0066] Multi-channel feature tensors can be determined based on the historical operating datasets of multiple lanes. The spatial neural network branch of the model can be determined using transmit power to extract features from the multi-channel feature tensors, resulting in multi-channel fusion features. These multi-channel fusion features can characterize the single-channel operating state of each lane and the inter-channel interference environment between lanes, thus effectively representing the complex mutual influence relationships between different lanes of a multimode optical module. For example, for lane 1, the multi-channel fusion features can characterize the single-channel operating state of lane 1 and the inter-channel interference environment of lane 1, and so on.

[0067] The actual transmit power of each lane at the current moment can be determined based on the multi-channel fusion characteristics and the candidate transmit power characteristics of each lane. For example, the actual transmit power of lane1 at the current moment can be determined based on the multi-channel fusion characteristics and the candidate transmit power characteristics of lane1, and so on.

[0068] According to embodiments of this application, a transmit power prediction model can be used to determine the transmit power of a multimode optical module at the current moment based on historical operating datasets of multiple lanes within the module. Specifically, the temporal neural network branch of the transmit power determination model can fully capture and learn the operational characteristics of each lane of the multimode optical module under conditions such as temperature changes and laser aging, based on the historical operating datasets of multiple lanes. The spatial neural network branch of the transmit power determination model can fully capture and learn the complex influence relationships and inter-channel environmental characteristics among the multiple lanes of the multimode optical module, based on the historical operating datasets of multiple lanes. The candidate transmit power features of each lane output by the temporal neural network branch can be fine-tuned based on the multi-channel fusion features output by the spatial neural network branch, thereby enabling the model to accurately predict the transmit power of each lane and effectively reducing the error between the predicted transmit power value and the actual transmit power value. Specifically:

[0069] The system can acquire historical operational datasets for each lane of a multimode optical module. The temporal neural network branch can analyze the variation patterns and dynamic trends of the transmit power of each lane over time based on the historical operational data sequences obtained from these datasets, outputting candidate transmit power features for each channel at the current moment. The spatial neural network branch can extract features from the multi-channel feature tensors determined by the historical operational datasets to deeply mine and capture the complex spatial correlations and mutual influences between different channels (lanes), obtaining multi-channel fusion features. Based on the multi-channel fusion features extracted by the spatial neural network branch, the candidate transmit power features for each lane output by the temporal neural network branch can be fine-tuned to address the mutual influence of power between different lanes, thereby obtaining the corrected actual predicted transmit power values ​​for each lane at the current moment (i.e., the true transmit power output by the model).

[0070] It is understood that, compared to using a fixed fitting function to fit the transmit power of an optical module, the transmit power determination method for multimode optical modules provided in this application can overcome the shortcomings of poor accuracy and low reliability in calculating transmit optical power by relying on the built-in fitting function of the optical module, and has the following technical effects:

[0071] 1. Eliminate errors introduced by temperature drift: Solve the problem of inaccurate estimated transmitted optical power caused by the failure of the factory calibration fitting function due to changes in the operating temperature of the optical module.

[0072] 2. Compensation for errors introduced by aging drift: This addresses the problem that the error in estimating emitted optical power continuously increases due to the drift of internal component parameters caused by the aging of optical modules after long-term use.

[0073] 3. Correcting errors caused by crosstalk between channels in multimode optical modules: This addresses the problem that the temperature drift error and component aging error are further amplified due to the mutual influence of power between multiple lanes in multimode optical modules.

[0074] 4. Improves the reliability of link status assessment: It fundamentally solves the problem of frequent false alarms and missed alarms in network management system link quality monitoring caused by the unreliability of the transmitted optical power data used as the basis for assessment, thereby improving the accuracy and efficiency of data center network operation and maintenance.

[0075] According to an embodiment of this application, historical operating data includes historical bias current, historical driving voltage, and historical temperature. Using a spatial neural network branch of the transmit power determination model, feature extraction is performed on the multi-channel feature tensor determined by multiple historical operating datasets to obtain multi-channel fusion features. This includes aligning and concatenating multiple historical bias currents, multiple historical driving voltages, and multiple historical temperatures from the historical operating datasets of each channel according to historical timestamps, forming single-channel feature matrices for each channel, and stacking these single-channel feature matrices according to channel dimensions to obtain a multi-channel feature tensor for multiple channels. Feature extraction is then performed on the multi-channel feature tensor based on the spatial neural network branch to obtain multi-channel fusion features.

[0076] In one embodiment, for any lane's historical operating dataset, the historical operating data may include historical bias current, historical drive voltage, and historical temperature. The historical bias current can be understood as the bias current of the laser in that lane at a historical moment, the historical drive voltage as the drive voltage of the multimode optical module at a historical moment, and the historical temperature as the operating temperature of the multimode optical module at a historical moment.

[0077] For any lane, multiple historical bias currents, multiple historical drive voltages, and multiple historical temperatures from the historical operating dataset can be aligned and concatenated according to historical timestamps to obtain a single-channel feature matrix for that lane. In one embodiment, the single-channel feature matrix is ​​a two-dimensional feature matrix, the dimensions of which include a time dimension and a parameter dimension, where the parameter dimension includes historical bias current, historical drive voltage, and historical temperature. Further, the single-channel feature matrices of multiple lanes can be stacked according to the channel dimension to obtain a multi-channel feature tensor for multiple lanes. For example, for lane 1, the single-channel feature matrix of lane 1 and the single-channel feature matrices of lanes 2, 3, and 4 related to lane 1 can be stacked according to the channel dimension to obtain a multi-channel feature tensor for lane 1. For example, for lane 2, the single-channel feature matrix of lane 2 and the single-channel feature matrices of lanes 1, 3, and 4 related to lane 2 can be stacked according to the channel dimension to obtain a multi-channel feature tensor for lane 2.

[0078] In one embodiment, the multi-channel feature tensor is a three-dimensional feature tensor, the dimensions of which include channel dimension, time dimension and parameter dimension, wherein the channel dimension is used to characterize multiple lanes of the optical module.

[0079] Feature extraction can be performed on multi-channel feature tensors based on spatial neural network branches to deeply mine and capture the complex spatial correlations and mutual influences between different lanes, obtaining multi-channel fusion features that characterize the single-channel operating state of multiple lanes and the inter-channel interference environment between multiple lanes. For example, for lane1, the multi-channel fusion features characterize the single-channel operating state of lane1 and the multi-channel fusion features of the inter-channel interference environment formed by lanes2, 3, and 4. For example, for lane2, the multi-channel fusion features characterize the single-channel operating state of lane2 and the multi-channel fusion features of the inter-channel interference environment formed by lanes1, 3, and 4.

[0080] According to an embodiment of this application, multiple historical bias currents, multiple historical driving voltages, and multiple historical temperatures from multiple channels in the historical operation dataset are aligned and concatenated according to historical timestamps to form single-channel feature matrices for each channel. These single-channel feature matrices are then stacked according to channel dimensions to obtain a multi-channel feature tensor for each channel. This includes: for each channel: arranging multiple historical bias currents sequentially according to historical timestamps to obtain a historical bias current sequence; arranging multiple historical driving voltages sequentially according to historical timestamps to obtain a historical driving voltage sequence; arranging multiple historical temperatures sequentially according to historical timestamps to obtain a historical temperature sequence; for each channel: aligning the historical bias current sequence, historical driving voltage sequence, and historical temperature sequence according to historical timestamps and concatenating them to obtain a two-dimensional feature matrix, which is then used as a single-channel feature matrix. The rows and columns of the two-dimensional feature matrix are arranged according to historical timestamps and a preset data order, respectively; stacking the single-channel feature matrices for each channel according to channel dimensions to obtain a three-dimensional feature tensor, which is then used as a multi-channel feature tensor.

[0081] Figure 3A A flowchart illustrating the construction of a multichannel feature tensor according to an embodiment of this application is shown.

[0082] like Figure 3A As shown, in operation S301, for each channel (lane), multiple historical driving voltages are arranged sequentially according to the order of historical timestamps (e.g., timestamp 1 to timestamp 3) to obtain a historical driving voltage sequence; multiple historical driving voltages are arranged sequentially according to the order of historical timestamps to obtain a historical driving voltage sequence; multiple historical temperatures are arranged sequentially according to the order of historical timestamps to obtain a historical temperature sequence.

[0083] In operation S302, for each channel, the historical bias current sequence, historical drive voltage sequence, and historical temperature sequence are concatenated after being timestamped according to the historical timestamp order to obtain a two-dimensional feature matrix, which is then used as the single-channel feature matrix. The rows and columns of the two-dimensional feature matrix can be arranged according to the historical timestamps and a preset data order, respectively. For example, for each channel, the rows of the two-dimensional feature matrix correspond to the time dimension and can be arranged in the chronological order of the historical timestamps, such as timestamp 1 to timestamp 3. For each channel, the columns of the two-dimensional feature matrix correspond to the parameter dimension, and can be arranged, for example, in the order of historical bias current, historical drive voltage, and historical temperature (i.e., the preset data order).

[0084] In operation S303, the single-channel feature matrices of multiple channels are stacked according to the channel dimension to obtain a three-dimensional feature tensor, and the three-dimensional feature tensor is used as a multi-channel feature tensor.

[0085] For example, for lane1, the two-dimensional feature matrices corresponding to lane1 and its related lanes 2, 3, and 4 (a total of 4 lanes) can be stacked according to channels to form a three-dimensional, multi-channel feature tensor with dimensions of 4 channels × 3 timestamps × 3 parameters. This three-dimensional feature tensor is then used as the multi-channel feature tensor corresponding to lane1. This multi-channel feature tensor can comprehensively represent the integrated state information of lane1 itself and the environment between its channels.

[0086] The multi-channel feature tensors of multiple lanes can be obtained in the above manner. The purpose of doing this is to combine the state parameters of different lanes under multiple modes to fully and comprehensively represent the common features of multiple lanes, so as to extract their high-dimensional feature information through spatial neural networks.

[0087] Figure 3B A schematic diagram of constructing a multichannel feature tensor according to an embodiment of this application is shown.

[0088] In one embodiment, the multimode optical module may include, for example, four channels: channel 1, channel 2, channel 3, and channel 4.

[0089] like Figure 3B As shown, the single-channel feature matrices for channels 1, 2, 3, and 4 are respectively feature matrix 1 (310), feature matrix 2 (320), feature matrix 3 (330), and feature matrix 4 (340). The single-channel feature matrix is ​​a two-dimensional feature matrix. For each channel, the rows of the two-dimensional feature matrix correspond to the time dimension, arranged in the order of timestamp 1 to timestamp 3, and the columns of the two-dimensional feature matrix correspond to the parameter dimension, arranged according to the historical bias current (…). ), historical driving voltage ( The order of ) and historical temperature (Temp).

[0090] like Figure 3B As shown, the single-channel feature matrices of multiple channels can be stacked according to the channel dimension to obtain a three-dimensional feature tensor 350. This three-dimensional feature tensor can then be used as a multi-channel feature tensor. For example, for channel 1, the first feature matrix 310 corresponding to channel 1, and the second feature matrix 320, third feature matrix 330, and fourth feature matrix 340 corresponding to channels 2, 3, and 4 respectively, can be stacked according to the channel dimension to form a three-dimensional, multi-channel feature tensor with dimensions of 4 channels × 3 timestamps × 3 parameters. This three-dimensional feature tensor can then be used as the multi-channel feature tensor 350 corresponding to channel 1.

[0091] According to embodiments of this application, the spatial neural network branch includes a spatial neural network branch based on a convolutional neural network.

[0092] Convolutional Neural Networks (CNNs) are a type of feedforward neural network that incorporates convolutional computations and has a deep structure. They are suitable for processing data with a grid-like structure. CNNs automatically extract features from the data by performing convolutional operations between the input data and the convolutional kernels in the convolutional layers. Compared to traditional neural networks, CNNs have several significant characteristics: local connectivity, weight sharing, and translation invariance. For example, the neurons in each layer of a CNN are arranged in three dimensions (width, height, and depth). For instance, for the input layer, width and height refer to the width and height of the input features, and depth represents the number of channels in the input features; for intermediate layers, width and height refer to the width and height of the feature map, usually determined by the parameters of the convolution and pooling operations, and depth refers to the number of channels in the feature map, usually determined by the number of convolutional kernels.

[0093] In related technologies, convolutional neural networks are widely used in image recognition and vision tasks. However, in this embodiment, since the multi-channel feature tensor is a three-dimensional feature tensor with dimensions of 4 channels × 3 timestamps × 3 parameters, it can be well adapted to the network structure of convolutional neural networks. This allows for the utilization of the characteristics of convolutional neural networks, such as "local receptive field, weight sharing, and translation invariance," to further extract high-dimensional features from the multi-channel feature tensor to achieve deep fusion of multi-channel information, resulting in multi-channel fused features. These multi-channel fused features can characterize the individual single-channel operating states of multiple channels and the inter-channel interference environment between multiple channels.

[0094] In one embodiment, for any lane, a constructed 4-channel × 3-timestamp × 3-parameter multi-channel feature tensor can be input into a spatial neural network branch. The core of this spatial neural network branch can be, for example, a 3×3 convolutional kernel or an equivalent spatial feature extractor. Through its internal hierarchical computation, the spatial neural network branch can perform spatial operations such as convolution and pooling on the input multi-channel feature tensor, aiming to deeply mine and capture the complex spatial correlations and mutual influences between different transmission lanes. The spatial neural network branch ultimately outputs a set of high-dimensional spatial feature vectors (i.e., multi-channel fused features). These multi-channel fused features condense key spatial pattern information under multi-channel collaborative working conditions, providing a data foundation for subsequent fusion and correction of candidate transmit power features output by the temporal neural network branch.

[0095] According to an embodiment of this application, the historical operation data also includes historical transmission power; for each channel, the time neural network branch of the transmission power determination model is used to determine the candidate transmission power features at the current moment based on the historical operation data sequence obtained from the historical operation data sequence, including: arranging multiple historical transmission powers in the historical operation dataset in the order of historical timestamps to obtain a historical transmission power sequence; inputting the historical transmission power sequence into the time neural network branch and outputting the candidate transmission power features at the current moment.

[0096] In one embodiment, for any lane's historical operating dataset, the historical operating data also includes historical transmit power, which can be understood as the optical transmit power of the lane at a historical moment.

[0097] Figure 4 A flowchart illustrating the determination of candidate transmit power characteristics for each of the multiple channels according to an embodiment of this application is shown.

[0098] like Figure 4 As shown, in operation S401, for each channel (lane), multiple historical transmit power values ​​from the historical running dataset are arranged sequentially according to their historical timestamps (e.g., timestamp 1 to timestamp 3) to obtain a historical transmit power sequence. For each lane, the historical transmit power sequence can reflect the historical trend of transmit power changes in that lane.

[0099] In operation S402, for each channel, the historical transmit power sequence is input into the temporal neural network branch, which outputs the candidate transmit power features at the current moment. The temporal neural network branch can analyze the change pattern and dynamic trend of transmit power over time based on the historical transmit power sequence.

[0100] For example, for lane1, the temporal neural network branch performs deep learning on the historical transmit power sequence of the input lane1, and finally outputs a preliminary prediction of the transmit power of lane1 at the current moment (that is, the candidate transmit power feature of lane1). This preliminary prediction is mainly based on the historical patterns of lane1 itself and is the basis for subsequent feature fusion and correction.

[0101] According to embodiments of this application, the temporal neural network branch includes at least one of the following: a temporal neural network branch based on a temporal convolutional network, a temporal neural network branch based on a recurrent neural network, and a temporal neural network branch based on a long short-term memory network.

[0102] Temporal Convolutional Networks (TCNs) are neural network architectures specifically designed for processing time-series data. They leverage the capabilities of convolutional layers to capture patterns in time series data, enabling efficient prediction and classification. The core idea of ​​TCNs is to use a one-dimensional convolutional network (1D CNN) to process sequential data while ensuring no information leakage occurs during convolution; that is, when predicting data at the current time step, the model can only use data from the current time step and earlier. As an example, a TCN consists of multiple convolutional layers, each containing causal convolutions and dilated convolutions. Within each convolutional layer, by choosing appropriate dilation factors and kernel sizes, the model can cover historical information of varying lengths. Furthermore, each convolutional layer in a TCN typically includes activation functions, normalization, and regularization operations to improve the model's generalization ability. In a TCN, the output at each time step is associated with previous time steps, and TCNs can process sequential data in parallel, thus improving efficiency during training and prediction.

[0103] Recurrent Neural Networks (RNNs) are neural networks specifically designed for processing sequential data, and are widely used in fields such as natural language processing, speech recognition, and time series prediction. The core of an RNN lies in its recurrent connection mechanism, where the output at the current time step depends not only on the current input but also on information from all previous time steps. By introducing hidden states, RNNs can capture temporal dependencies within a sequence. Through the propagation of hidden states, RNNs achieve dynamic modeling of sequential data.

[0104] Long Short-Term Memory Networks (LSTM) are a type of temporal recurrent neural network suitable for processing and predicting important events with relatively long intervals and delays in time series. LSTM was proposed to address the "vanishing gradient" problem in RNN structures and is a special type of recurrent neural network. LSTM effectively alleviates the vanishing gradient problem by introducing input gates, forget gates, and output gates to selectively remember or discard information.

[0105] According to an embodiment of this application, the transmit power determination model further includes an output module; determining the true transmit power of each of the multiple channels at the current moment based on the multi-channel fusion features and the candidate transmit power features of each of the multiple channels includes: based on the fusion submodule of the output module, fusing the multi-channel fusion features with the candidate transmit power features of each of the multiple channels respectively to obtain the target fusion features after correcting the candidate transmit power features of each of the multiple channels; the target fusion features characterize the transmit power change trend of the channel and the inter-channel interference environment in which the channel is located; based on the prediction submodule of the output module, performing regression prediction on the target fusion features of each of the multiple channels respectively, and outputting the true transmit power of each of the multiple channels.

[0106] In one embodiment, the transmit power also includes an output module. The output module may include a fusion submodule and a prediction submodule.

[0107] For any lane, the multi-channel fusion features for that lane can be fused with the candidate transmit power features of that lane using the fusion submodule, resulting in the target fusion features after correcting the candidate transmit power features of that lane. For example, for lane 1, the multi-channel fusion features for lane 1 (output by the spatial neural network branch) can be fused with the candidate transmit power features of lane 1 (output by the temporal neural network branch) to obtain the target fusion features of lane 1. It can be understood that the target fusion features of lane 1 simultaneously include the high-dimensional spatial features extracted from the spatial neural network branch (reflecting the mutual influence between multiple channels) and the preliminary predicted values ​​output from the temporal neural network branch (reflecting the power change trend of lane 1 itself).

[0108] For any lane, the target fusion features of the lane can be input into the output module based on the prediction submodule to perform regression prediction and obtain the true transmit power of the lane.

[0109] According to an embodiment of this application, the prediction submodule includes a fully connected neural network. Based on the prediction submodule, regression prediction is performed on the target fusion features of each of the multiple channels respectively, and the corresponding output of the true transmission power of each of the multiple channels includes: for each channel: the target fusion features are nonlinearly transformed based on the fully connected neural network to obtain hidden features, and regression prediction is performed on the hidden features to output the true transmission power.

[0110] Figure 5 A schematic diagram of a fully connected neural network according to an embodiment of this application is shown.

[0111] In one embodiment, a fully connected neural network includes an input layer L510 (INPUT), at least one hidden layer L520 (HIDDEN), and an output layer L530 (OUTPUT).

[0112] For any lane, such as Figure 5 As shown, the input layer receives the target fused features. The hidden layer, composed of multiple neurons, performs deeper nonlinear transformations and feature abstraction on the target fused features to learn an accurate feature representation for the final power prediction, thus obtaining the hidden features. After computation in the hidden layer, the hidden features are passed to the output layer. The output layer consists of a single neuron that performs the final linear or nonlinear transformation. Finally, the output layer generates and outputs a specific numerical value as the true transmit power, i.e., the "predicted true transmit power of lane at the current moment".

[0113] Figure 6 A schematic diagram of a method for determining the transmit power of a multimode optical module according to an embodiment of this application is shown.

[0114] like Figure 6 As shown, by operating S601, the historical operating datasets of each of the multiple channels of the multimode optical module can be obtained.

[0115] When operating S602, a historical transmit power sequence can be constructed based on multiple historical transmit powers in the historical operation datasets of multiple channels.

[0116] When operating the S603, the time neural network branch of the transmit power determination model can be used to determine the candidate transmit power characteristics of each of the multiple channels at the current moment based on the historical operating data sequences of each of the multiple channels.

[0117] When operating the S604, a multi-channel feature tensor can be constructed based on multiple historical bias currents, multiple historical drive voltages, and multiple historical temperatures from the historical operating datasets of each of the multiple channels.

[0118] When operating the S605, the spatial neural network branch of the model can be determined by the transmit power, and features can be extracted from the multi-channel feature tensor to obtain multi-channel fused features.

[0119] When operating S606, the multi-channel fusion feature can be fused with the candidate transmit power features of each of the multiple channels to obtain the target fusion feature after correcting the candidate transmit power features of each of the multiple channels.

[0120] When operating the S607, a fully connected neural network can be used to perform regression prediction on the target fusion features of multiple channels respectively, and output the true transmit power of each channel.

[0121] Figure 7 A flowchart is shown of a training method for a transmit power determination model for a multimode optical module according to an embodiment of this application.

[0122] like Figure 7 As shown, the training method for the transmit power determination model of the multimode optical module includes operations S710 to S750.

[0123] When operating the S710, obtain the sample running datasets for each of the multiple sample channels of the sample multimode optical module.

[0124] The sample run dataset includes multiple sample run data with sample timestamps.

[0125] When operating the S720, the time neural network branch of the candidate transmit power determination model is used to determine the first transmit power feature of each of the multiple sample channels at the current time based on the sample running data sequence obtained from the sample running datasets of each of the multiple sample channels.

[0126] When operating the S730, the spatial neural network branch of the candidate transmit power determination model is used to extract features from the multi-channel feature tensor of the samples determined by multiple sample running datasets, and the multi-channel fused features of the samples are obtained.

[0127] The multi-channel fusion feature of the sample represents the single-channel operating status of multiple sample channels and the inter-channel interference environment between the multi-sample channel environments.

[0128] In operation of S740, based on the multi-channel fusion characteristics of the samples and the first transmit power characteristics of each of the multiple sample channels, the predicted transmit power values ​​of each of the multiple sample channels at the current time are determined.

[0129] When operating the S750, the candidate transmit power determination model is trained based on the predicted transmit power values ​​and the actual transmit power values ​​of multiple sample channels to obtain the transmit power determination model.

[0130] The true transmit power values ​​of each of the multiple sample channels are: the optical power received by the peer optical module that transmits data with the sample multimode optical module through the multiple sample channels at the current moment of the sample.

[0131] In one embodiment, the candidate transmit power determination model can be trained based on the sample running datasets of multiple sample lanes of the sample multimode optical module to obtain the transmit power determination model. For any sample lane, the sample running dataset includes multiple sample running data with sample timestamps. The sample timestamp corresponds to the sample time, and the sample running data can be understood as the running status data at the sample time.

[0132] For example, the candidate transmit power determination model may include a temporal neural network branch and a spatial neural network branch. Based on the sample run datasets of multiple sample lanes, sample run data sequences for each of the multiple sample lanes can be obtained. Using the temporal neural network branch of the candidate transmit power determination model, the first transmit power feature of each of the multiple sample lanes at the current time can be determined based on the sample run data sequences of each of the multiple sample lanes. The multi-channel feature tensor of the sample can be determined based on the sample run datasets of each of the multiple sample lanes. Using the spatial neural network branch of the candidate transmit power determination model, feature extraction can be performed on the multi-channel feature tensor to obtain the multi-channel fused feature of the sample. Based on the multi-channel fused feature and the first transmit power feature of each of the multiple sample lanes, the predicted transmit power value of each of the multiple sample lanes at the current time can be determined.

[0133] The candidate transmit power determination model can be trained based on the predicted transmit power values ​​and the actual transmit power values ​​of multiple sample lanes until a predetermined termination condition is met, thus obtaining the transmit power determination model. The predetermined termination condition can be, for example, reaching a predetermined number of training iterations or the convergence of the loss function.

[0134] The true transmit power of each of the multiple sample lanes is: the optical power received by the peer optical module transmitting data with the sample multimode optical module through the multiple sample lanes at the current moment. The optical power received by the peer optical module through the multiple sample lanes can be understood as the independent received optical power measured by the peer optical module on each sample lane, which serves as the true label of the transmit power of the corresponding sample lane.

[0135] In one embodiment, new operating status data of the multimode optical module can be collected at a predetermined period, and the transmission power determination model can be fine-tuned and trained based on the new operating status data so that the transmission power determination model can dynamically adapt to various operating conditions and changing trends of the multimode optical module, ensuring the accuracy and reliability of the model.

[0136] Figure 8 The diagram schematically illustrates a structural block diagram of a transmit power determination device for a multimode optical module according to an embodiment of this application.

[0137] like Figure 8 As shown, the transmit power determination device for a multimode optical module includes a first acquisition module 810, a first processing module 820, a second processing module 830, and a first determination module 840.

[0138] The first acquisition module 810 is used to acquire the historical operation datasets of each of the multiple channels of the multimode optical module; the historical operation datasets include multiple historical operation data with historical timestamps.

[0139] The first processing module 820 is used to determine the candidate transmission power characteristics of multiple channels at the current moment by utilizing the time neural network branch of the transmission power determination model and based on the historical operation data sequence obtained from the historical operation datasets of multiple channels.

[0140] The second processing module 830 is used to determine the spatial neural network branch of the transmission power model, extract features from the multi-channel feature tensor determined by multiple historical operation datasets, and obtain multi-channel fusion features; the multi-channel fusion features characterize the single-channel operation status of multiple channels and the inter-channel interference environment between multiple channels.

[0141] The first determining module 840 is used to determine the actual transmit power of each of the multiple channels at the current moment based on the multi-channel fusion characteristics and the candidate transmit power characteristics of each of the multiple channels.

[0142] According to embodiments of this application, historical operating data includes historical bias current, historical drive voltage, and historical temperature. The second processing module 830 may include a first processing submodule and a second processing submodule.

[0143] The first processing submodule is used to align and stitch together multiple historical bias currents, multiple historical driving voltages and multiple historical temperatures in the historical running datasets of multiple channels according to historical timestamps to form single-channel feature matrices for each channel. The single-channel feature matrices of multiple channels are then stacked according to the channel dimension to obtain a multi-channel feature tensor for multiple channels.

[0144] The second processing submodule is used to extract features from the multi-channel feature tensor based on the spatial neural network branch to obtain multi-channel fused features.

[0145] According to embodiments of this application, the first processing submodule may include a first processing unit, a second processing unit, and a third processing unit.

[0146] The first processing unit is used for each channel to: arrange multiple historical bias currents in the order of historical timestamps to obtain a historical bias current sequence; arrange multiple historical drive voltages in the order of historical timestamps to obtain a historical drive voltage sequence; and arrange multiple historical temperatures in the order of historical timestamps to obtain a historical temperature sequence.

[0147] The second processing unit is used for each channel to: align the historical bias current sequence, historical drive voltage sequence, and historical temperature sequence with historical timestamps and then concatenate them to obtain a two-dimensional feature matrix. The two-dimensional feature matrix is ​​then used as a single-channel feature matrix, with the rows and columns of the two-dimensional feature matrix arranged according to the historical timestamps and the preset data order, respectively.

[0148] The third processing unit is used to stack the single-channel feature matrices of multiple channels according to the channel dimension to obtain a three-dimensional feature tensor, and use the three-dimensional feature tensor as a multi-channel feature tensor.

[0149] According to embodiments of this application, the historical operating data also includes historical transmission power. The first processing module 820 may include a fourth processing unit and a fifth processing unit.

[0150] The fourth processing unit is used to arrange multiple historical transmission powers in the historical operation dataset in the order of historical timestamps to obtain a historical transmission power sequence.

[0151] The fifth processing unit is used to input the historical transmit power sequence into the time neural network branch and output the candidate transmit power features at the current moment.

[0152] According to embodiments of this application, the transmit power determination model further includes an output module. The first determination module 840 may include a sixth processing unit and a seventh processing unit.

[0153] The sixth processing unit is used to fuse the multi-channel fusion features with the candidate transmit power features of each of the multiple channels based on the fusion submodule of the output module, so as to obtain the target fusion features after correcting the candidate transmit power features of each of the multiple channels; the target fusion features characterize the transmit power change trend of the channel and the inter-channel interference environment in which the channel is located.

[0154] The seventh processing unit is used to perform regression prediction on the target fusion features of each of the multiple channels based on the prediction submodule of the output module, and output the true transmit power of each of the multiple channels.

[0155] According to an embodiment of this application, the prediction submodule includes a fully connected neural network. The seventh processing unit is used to: for each channel, perform a nonlinear transformation on the target fusion features based on the fully connected neural network to obtain hidden features; and perform regression prediction on the hidden features to output the true transmit power.

[0156] According to embodiments of this application, the temporal neural network branch includes at least one of the following: a temporal neural network branch based on a temporal convolutional network, a temporal neural network branch based on a recurrent neural network, and a temporal neural network branch based on a long short-term memory network.

[0157] According to embodiments of this application, the spatial neural network branch includes a spatial neural network branch based on a convolutional neural network.

[0158] According to embodiments of this application, any plurality of modules among the first acquisition module 810, the first processing module 820, the second processing module 830, and the first determination module 840 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in one module. According to embodiments of this application, at least one of the first acquisition module 810, the first processing module 820, the second processing module 830, and the first determination module 840 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in any one of the three implementation methods of software, hardware, and firmware, or in a suitable combination of any of these. Alternatively, at least one of the first acquisition module 810, the first processing module 820, the second processing module 830, and the first determination module 840 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.

[0159] Figure 9 The diagram schematically illustrates a structural block diagram of a training apparatus for determining the transmit power of a multimode optical module according to an embodiment of this application.

[0160] like Figure 9 As shown, the training device for determining the transmit power model of a multimode optical module includes a second acquisition module 910, a third processing module 920, a fourth processing module 930, a second determination module 940, and a training module 950.

[0161] The second acquisition module 910 acquires the sample running datasets of each of the multiple sample channels of the sample multimode optical module; the sample running datasets include multiple sample running data with sample timestamps.

[0162] The third processing module 920 uses the time neural network branch of the candidate transmit power determination model to determine the first transmit power feature of each of the multiple sample channels at the current time based on the sample running data sequence obtained from the sample running datasets of each of the multiple sample channels.

[0163] The fourth processing module 930 uses the spatial neural network branch of the candidate transmit power determination model to extract features from the sample multi-channel feature tensor determined by multiple sample running datasets, and obtains sample multi-channel fusion features; the sample multi-channel fusion features characterize the single-channel running state of multiple sample channels and the inter-channel mutual interference environment between multiple sample channel environments.

[0164] The second determining module 940 determines the predicted transmission power of each of the multiple sample channels at the current moment based on the multi-channel fusion characteristics of the samples and the first transmission power characteristics of each of the multiple sample channels.

[0165] Training module 950 is used to train the candidate transmission power determination model based on the predicted transmission power values ​​and the actual transmission power values ​​of the multiple sample channels to obtain the transmission power determination model; wherein, the actual transmission power values ​​of the multiple sample channels are: the optical power received by the peer optical module that transmits data with the sample multimode optical module through the multiple sample channels at the current time of the sample.

[0166] According to embodiments of this application, any multiple modules among the second acquisition module 910, the third processing module 920, the fourth processing module 930, the second determination module 940, and the training module 950 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module. According to embodiments of this application, at least one of the second acquisition module 910, the third processing module 920, the fourth processing module 930, the second determination module 940, and the training module 950 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the second acquisition module 910, the third processing module 920, the fourth processing module 930, the second determination module 940, and the training module 950 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.

[0167] Figure 10 The diagram schematically illustrates an electronic device suitable for implementing a method for determining the transmit power of a multimode optical module and a method for training a model for determining the transmit power of a multimode optical module, according to embodiments of this application.

[0168] like Figure 10 As shown, an electronic device 1000 according to an embodiment of this application includes a processor 1001, which can perform various appropriate actions and processes according to a program stored in a ROM 1002 (read-only memory) or a program loaded from a storage portion 1008 into a RAM 1003 (random access memory). The processor 1001 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 1001 may also include onboard memory for caching purposes. The processor 1001 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.

[0169] RAM 1003 stores various programs and data required for the operation of electronic device 1000. Processor 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Processor 1001 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 1002 and / or RAM 1003. It should be noted that the programs may also be stored in one or more memories other than ROM 1002 and RAM 1003. Processor 1001 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in said one or more memories.

[0170] According to embodiments of this application, the electronic device 1000 may further include an input / output (I / O) interface 1005, which is also connected to a bus 1004. The electronic device 1000 may also include one or more of the following components connected to the input / output (I / O) interface 1005: an input section 1006 including a keyboard, mouse, etc.; an output section 1007 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN card, modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the input / output (I / O) interface 1005 as needed. A removable medium 1011, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1010 as needed so that computer programs read from it can be installed into the storage section 1008 as needed.

[0171] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the methods described in the embodiments of this application.

[0172] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 1002 and / or RAM 1003 and / or one or more memories other than ROM 1002 and RAM 1003 described above.

[0173] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to cause the computer system to implement the methods provided in the embodiments of this application.

[0174] When the computer program is executed by the processor 1001, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0175] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 1009, and / or installed from a removable medium 1011. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0176] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009, and / or installed from the removable medium 1011. When the computer program is executed by the processor 1001, it performs the functions defined in the system of this application embodiment. According to the embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0177] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0178] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0179] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.

[0180] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.

Claims

1. A method for determining the transmit power of a multimode optical module, characterized in that, A multimode optical module is used for optical communication in data center networks. The multimode optical module includes multiple parallel channels for data transmission. The method includes: Obtain the historical operation datasets for each of the multiple channels of the multimode optical module; the historical operation datasets include multiple historical operation data with historical timestamps; The time neural network branch of the transmit power determination model is used to determine the candidate transmit power characteristics of each of the multiple channels at the current moment based on the historical operation data sequence obtained from the historical operation datasets of each of the multiple channels. By utilizing the spatial neural network branch of the transmit power determination model, feature extraction is performed on the multi-channel feature tensor determined by multiple historical operation datasets to obtain multi-channel fusion features. The multi-channel fusion features characterize the single-channel operation status of each of the multiple channels and the inter-channel interference environment between the multiple channels. The fusion submodule of the output module of the transmit power determination model is used to fuse the multi-channel fusion features with the candidate transmit power features of each of the multiple channels respectively, so as to obtain the target fusion features after correcting the candidate transmit power features of each of the multiple channels; the target fusion features characterize the transmit power variation trend of the channel and the inter-channel interference environment in which the channel is located; The prediction submodule of the output module of the model is determined by using the transmission power to perform regression prediction on the target fusion features of each of the multiple channels, and output the true transmission power of each of the multiple channels.

2. The method according to claim 1, characterized in that, Historical operating data includes historical bias current, historical drive voltage, and historical temperature; using the spatial neural network branch of the transmit power determination model, feature extraction is performed on the multi-channel feature tensor determined by multiple historical operating datasets to obtain multi-channel fused features, including: The historical bias currents, historical driving voltages, and historical temperatures from the historical operation datasets of each channel are aligned and concatenated according to historical timestamps to form single-channel feature matrices for each channel. The single-channel feature matrices of each channel are then stacked according to the channel dimension to obtain a multi-channel feature tensor for each channel. Based on the spatial neural network branch, feature extraction is performed on the multi-channel feature tensor to obtain the multi-channel fused feature.

3. The method according to claim 2, characterized in that, The historical bias currents, historical drive voltages, and historical temperatures from the historical operating datasets of each channel are aligned and concatenated according to historical timestamps to form single-channel feature matrices for each channel. These single-channel feature matrices are then stacked according to channel dimensions to obtain a multi-channel feature tensor for each channel, including: For each channel: arrange multiple historical bias currents in order of their historical timestamps to obtain a historical bias current sequence; arrange multiple historical drive voltages in order of their historical timestamps to obtain a historical drive voltage sequence; arrange multiple historical temperatures in order of their historical timestamps to obtain a historical temperature sequence. For each channel: according to the order of historical timestamps, the historical bias current sequence, historical drive voltage sequence and historical temperature sequence are timestamp aligned and then concatenated to obtain a two-dimensional feature matrix. The two-dimensional feature matrix is ​​then used as the single-channel feature matrix, and the rows and columns of the two-dimensional feature matrix are arranged according to the historical timestamps and the preset data order, respectively. The single-channel feature matrices of multiple channels are stacked according to the channel dimension to obtain a three-dimensional feature tensor, and the three-dimensional feature tensor is used as the multi-channel feature tensor.

4. The method according to claim 1, characterized in that, Historical operational data also includes historical transmit power; for each channel, the time neural network branch of the transmit power determination model is used to determine the candidate transmit power features at the current moment based on the historical operational data sequence. The historical transmission power sequence is obtained by arranging multiple historical transmission powers in the historical runtime dataset in the order of historical timestamps. The historical transmit power sequence is input into the time neural network branch, and the candidate transmit power features at the current moment are output.

5. The method according to claim 1, characterized in that, The prediction submodule includes a fully connected neural network. Based on the output module, the prediction submodule performs regression prediction on the target fusion features of each of the multiple channels, and outputs the true transmit power of each of the multiple channels, including, for each channel: Based on a fully connected neural network, the target fusion features are subjected to a nonlinear transformation to obtain hidden features; It then performs regression prediction on the hidden features and outputs the true transmission power.

6. The method according to claim 1 or 4, characterized in that, The temporal neural network branch includes at least one of the following: a temporal neural network branch based on temporal convolutional networks, a temporal neural network branch based on recurrent neural networks, and a temporal neural network branch based on long short-term memory networks.

7. The method according to claim 1 or 2, characterized in that, The spatial neural network branch includes the spatial neural network branch based on convolutional neural networks.

8. A training method for a transmit power determination model of a multimode optical module, characterized in that, The method includes: Obtain the sample running datasets for each of the multiple sample channels of the sample multimode optical module; the sample running datasets include multiple sample running data with sample timestamps; The time neural network branch of the candidate transmit power determination model is used to determine the first transmit power feature of each of the multiple sample channels at the current time based on the sample running data sequence obtained from the sample running datasets of each of the multiple sample channels. By utilizing the spatial neural network branch of the candidate transmit power determination model, feature extraction is performed on the sample multi-channel feature tensor determined by multiple sample running datasets to obtain sample multi-channel fusion features. The sample multi-channel fusion features characterize the single-channel running state of multiple sample channels and the inter-channel mutual interference environment between multiple sample channel environments. The fusion submodule of the output module of the candidate transmit power determination model is used to fuse the multi-channel fusion features of the samples with the first transmit power features of each of the multiple sample channels respectively, so as to obtain the features after correcting the first transmit power features of each of the multiple sample channels; the corrected features characterize the transmit power change trend of the sample channel and the inter-channel interference environment in which the sample channel is located. The prediction submodule of the output module of the candidate transmit power determination model is used to perform regression prediction on the corrected features of each of the multiple sample channels, and output the transmit power prediction values ​​of each of the multiple channels. The candidate transmit power determination model is trained based on the predicted transmit power values ​​and the actual transmit power values ​​of multiple sample channels to obtain the transmit power determination model. The actual transmit power values ​​of multiple sample channels are the optical power received by the peer optical module that transmits data with the sample multimode optical module through multiple sample channels at the current time of the sample.

9. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Data prediction method, device and equipment

    CN120263282A

  • Transmitting power determination method and device, electronic equipment, medium and product

    CN121325127A