Starting method for pre-starting execution environment, computer equipment and related products

By introducing an independent monitoring timer and retransmission mechanism into the pre-boot execution environment, the deadlock problem caused by task priority competition between network drivers is solved, and the system's self-recovery and the reliability of the startup process are improved.

CN121050784AActive Publication Date: 2025-12-02INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202511558282.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-29
Publication Date
2025-12-02
Estimated Expiration
2045-10-29

AI Technical Summary

Technical Problem

In the pre-boot execution environment, deadlock caused by task priority competition between the first network driver and the second network driver affects the reliability and efficiency of the system startup process.

Method used

A monitoring timer independent of the target event polling operation is introduced. By monitoring the timeout event and network reply message event of the timer, the first network driver is triggered to resend the network release message. After the number of retransmissions exceeds the threshold, the target event polling operation is forcibly terminated, and network control is transferred to the second network driver.

Benefits of technology

It resolves the polling blocking problem caused by task priority contention, ensures that the system can recover itself in the event of a deadlock, and improves the system's ability to detect abnormal states and the reliability of the startup process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121050784A_ABST
    Figure CN121050784A_ABST
Patent Text Reader

Abstract

The invention discloses a starting method for a pre-starting execution environment, computer equipment and a related product, and relates to the technical field of computers, the method comprises the following steps: when a first network driver triggers a target unloading process due to loading of a second network driver and the first network driver has sent a network release message in the target unloading process, starting the target unloading process; monitoring an overtime event and a network reply message event of the monitoring timer; when the timeout event is monitored and the network reply message event is not monitored, triggering the first network driver to resend the network release message; and when the retransmission times exceed a preset retransmission threshold value, forcibly terminating the target event polling operation, and transferring the network control right from the first network driver to the second network driver, so that the second network driver takes over and continues to execute the starting process based on the own network protocol stack. The technical problem of deadlock in the starting process of the pre-starting execution environment in related technologies is solved, and the technical effect of improving the starting reliability of the pre-starting execution environment is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method for starting a pre-boot execution environment, computer equipment, and related products. Background Technology

[0002] In the pre-boot execution environment, when the system completes the loading of the network boot program and transfers control from the firmware's native first network driver to the high-performance second network driver, a network driver unloading process needs to be executed. During this process, the first network driver needs to send a network release message to the server and enter a target event polling operation to wait for the corresponding network reply message, thereby confirming the successful release of network resources.

[0003] However, in practical applications, it has been found that because the loading of the second network driver triggers dynamic adjustments to system task priorities, the first network driver is highly susceptible to task priority contention with the second network driver when performing target event polling operations. This contention causes the polling channel of the first network driver to be blocked, making it unable to effectively capture network response messages. Even if the server side has sent messages normally, the first network driver will be blocked indefinitely in the target event polling operation, unable to successfully complete the unloading process and transfer network control. This can lead to a deadlock in the system startup process, preventing the pre-startup execution environment from being established normally, and consequently affecting the reliability and efficiency of automated deployment of the server cluster. Summary of the Invention

[0004] This application provides a startup method, computer device, and related products for a pre-boot execution environment, to solve the technical problem in the related art where, during the transfer of network driver control during the startup of a pre-boot execution environment, a task priority conflict occurs between the first network driver and the second network driver, causing the system to remain in a state of waiting for network response messages, ultimately leading to a deadlock.

[0005] This application provides a method for starting a pre-boot execution environment, the method comprising: In response to the first network driver triggering the target unloading process due to the loading of the second network driver, and the first network driver having already sent a network release message during the target unloading process, a monitoring timer independent of the target event polling operation is started. The target event polling operation is the event polling operation in which the first network driver waits for the network reply message corresponding to the network release message to be executed. In response to the start of the monitoring timer, the timeout event of the monitoring timer and the network reply message event are monitored simultaneously. In response to the monitoring of a timeout event but no network reply message event, the first network driver is triggered to resend the network release message. In response to the number of times the network release message is resent exceeding the preset retransmission threshold, the target event polling operation is forcibly terminated, and network control is transferred from the first network driver to the second network driver, so that the second network driver can take over and continue to execute the startup process of the pre-boot execution environment based on its own network protocol stack.

[0006] This application also provides a computer device, including: a memory for storing a computer program; and a processor for implementing the startup method of the pre-boot execution environment in the following embodiments when executing the computer program.

[0007] In response to the first network driver triggering the target unloading process due to the loading of the second network driver, and the first network driver having already sent a network release message during the target unloading process, a monitoring timer independent of the target event polling operation is started. The target event polling operation is the event polling operation in which the first network driver waits for the network reply message corresponding to the network release message to be executed. In response to the start of the monitoring timer, the timeout event of the monitoring timer and the network reply message event are monitored simultaneously. In response to the monitoring of a timeout event but no network reply message event, the first network driver is triggered to resend the network release message. In response to the number of times the network release message is resent exceeding the preset retransmission threshold, the target event polling operation is forcibly terminated, and network control is transferred from the first network driver to the second network driver, so that the second network driver can take over and continue to execute the startup process of the pre-boot execution environment based on its own network protocol stack.

[0008] The pre-boot execution environment startup method provided in this application establishes an independent and undisturbed monitoring path by creating a monitoring timer independent of the original event polling operation. This addresses the polling blocking problem caused by task priority contention, ensuring that even if the main event polling loop of the first network driver is stuck due to resource contention, the timer timeout event can still be triggered, providing the possibility for system self-recovery. By setting a dual-channel monitoring mode, simultaneously monitoring both the timer timeout event and the network reply message event, the system can detect and react immediately regardless of which channel triggers the event first, improving the system's ability to perceive abnormal states. Setting the system to actively trigger the first network driver to resend the network release message when a timeout event is detected but no network reply message event is detected effectively addresses reply loss issues caused by momentary network packet loss or incorrect packet discarding, giving the system preliminary self-repair capabilities. By setting the number of times the network release message is resent to exceed a preset retransmission threshold, the target event polling operation is forcibly terminated, and network control is transferred from the first network driver to the second network driver. In this way, regardless of whether the release process is successful or not, the system can transfer network control to the second network driver, so that the second network driver can take over and continue to execute the startup process of the pre-boot execution environment based on its own network protocol stack, thus breaking the deadlock state. Attached Figure Description

[0009] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 A flowchart illustrating a startup method for a pre-boot execution environment provided in an embodiment of this application; Figure 2 A flowchart illustrating a startup method for a pre-boot execution environment provided in another embodiment of this application; Figure 3 A schematic diagram of a pre-boot execution environment provided in an embodiment of this application; Figure 4 A schematic diagram of the structure of a server providing a pre-boot execution environment according to an embodiment of this application; Figure 5 A flowchart illustrating the startup method of a pre-boot execution environment provided in another embodiment of this application; Figure 6 This is a schematic diagram of the structure of a startup device for a pre-startup execution environment provided in an embodiment of this application; Figure 7This is an internal structural diagram of a computer device provided in an embodiment of this application. Detailed Implementation

[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0012] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0013] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0014] like Figure 1 As shown, one embodiment of this application provides a method for starting a pre-boot execution environment, which specifically includes the following steps: Step 101: In response to the first network driver triggering the target unloading process due to the loading of the second network driver, and the first network driver having sent a network release message in the target unloading process, start a monitoring timer independent of the target event polling operation. The target event polling operation is the event polling operation executed by the first network driver waiting for the network reply message corresponding to the network release message.

[0015] The first network driver refers to the firmware's native network protocol stack driver, specifically the DHCPv6 network driver (Dynamic Host Configuration Protocol version 6 driver) embedded in the UEFI BIOS (based on the Unified Extensible Firmware Interface Basic Input / Output System). This driver is responsible for obtaining the network address and boot information during the initial startup phase.

[0016] The second network driver refers to a high-performance or feature-enhanced network driver that is loaded subsequently. Specifically, it can be the network driver of a network bootloader such as iPXE (an open-source preboot execution environment). This driver is designed to take over the hardware to provide faster network speeds or support more complex protocols.

[0017] The target unloading process refers to a series of orderly operations during the handover of control of the pre-startup execution environment, in which the first network driver is notified by the system to exit and needs to perform resource release and state cleanup.

[0018] The target event polling operation refers to a waiting mechanism in which the first network driver, after sending a release message, raises the task priority to the TPL_CALLBACK (callback-level network task priority level) level and cyclically checks whether data has arrived at the network interface.

[0019] A network release message is a protocol message initiated by the client to release previously allocated network resources. It contains information such as the IPv6 address to be released and the device identifier.

[0020] A network response message is an acknowledgment message returned by the server to the client after receiving a release message, used to confirm that the resource has been released.

[0021] The monitoring timer can be a software timer created based on UEFI timer events, independent of the network polling loop. The timer's timeout event has the same priority as the network event, ensuring that the timeout event can be responded to even if network polling is blocked, thus providing an independent time monitoring channel for the entire process.

[0022] It is understood that the first network driver in this application can refer to any software network stack entity responsible for initial network configuration and communication during the pre-boot or early boot phase, and which will be replaced or unloaded later in the boot process. Specifically, any driver that needs to be replaced by another network driver before the operating system kernel takes over completely falls under the category of the first network driver. The second network driver can refer to any software network stack entity loaded later in the boot process, which aims to take over or enhance network functions and thus triggers the unloading of the previous driver. Specifically, any driver loaded during the boot process that needs to exclusively or primarily control the network hardware falls under the category of the second network driver. At the same time, deadlock in this application can also refer to any stalemate state in which the system process cannot continue to advance due to irreconcilable competition between two or more drivers (or system components) over task priority, interrupt requests, memory access, or hardware register control. Specifically, any deadlock problem that can be solved by this application is caused by resource contention preventing the driver uninstallation process from being completed, thereby blocking the entire system boot sequence. This application only describes the case where the first network driver is the DHCPv6 network driver embedded in UEFIBIOS and the second network driver is iPXE, but it is not limited to this.

[0023] For example, suppose the deadlock is a startup failure caused by network quality issues. Specifically, in an unstable network environment with network congestion, high packet loss rate, or high latency, network reply messages / network release messages may be lost, causing the first network driver to wait indefinitely and fail to start due to timeout. The monitoring timer timeout and intelligent retransmission mechanism in this application can automatically initiate retransmission after a packet is lost, effectively combating network jitter and significantly improving the startup success rate in poor network environments.

[0024] For example, suppose the deadlock is a startup delay caused by slow server response. Specifically, when the PXE (Preboot Execution Environment) server is overloaded, it may be unable to process and respond to the client's release request in a timely manner, causing the client to wait for a long time, greatly extending the entire startup cycle. This application uses a pre-defined retransmission threshold and forced exit mechanism to set a maximum tolerable time for the entire waiting process. This avoids the entire startup process being suspended indefinitely due to the unresponsiveness of a single server, ensuring the predictability of startup time.

[0025] In one embodiment, during the initialization of the target unloading process, a dedicated storage area is configured to persistently store the protocol context of the current network session. The network session is the entire process of the first network driver sending a network release message and receiving a network reply message corresponding to the network release message. The protocol context includes key parameters used to uniquely identify this session, the client device, and the network resources to be released. The key parameters include a transaction identifier and a device unique identifier. The firmware interface is called to create a monitoring timer based on a hardware clock source. In response to the completion of the network release message transmission, the monitoring timer is initialized and started, and the timeout period of the monitoring timer is set to the initial threshold. The interrupt response of the monitoring timer is not affected by the task priority level of the first network driver.

[0026] A dedicated storage area can be a protected area in system memory used to securely save critical data during turbulent periods of driver unloading and switching, ensuring that it is not overwritten or lost by subsequent loaded drivers or cleanup processes.

[0027] A transaction identifier can be a unique random number that is unique within a complete network conversation. It is used to precisely match the client's request with the server's response and prevent message confusion.

[0028] A device unique identifier is a number that uniquely identifies a client device in the network, ensuring that the release message specifies which device's resources are to be released.

[0029] At the start of the target unloading process, the system configures a dedicated storage area in memory to persistently store the protocol context of the current network session. This session covers the entire process from sending a network release message to receiving a network reply message, and its protocol context includes key parameters such as transaction identifiers and device unique identifiers to uniquely identify the session, the client, and the resources to be released. Next, the system calls the firmware interface to create a monitoring timer based on a hardware clock source. Once the network release message is sent, the timer is immediately initialized and started, with its timeout set to an initial threshold. By configuring the interrupt response of this monitoring timer to be unaffected by the task priority level of the first network driver, it ensures that even if the main driver process is blocked due to resource contention, the timer timeout event can still be reliably triggered, providing an independent monitoring channel for system recovery from deadlock.

[0030] Step 102: In response to the start of the monitoring timer, monitor the timeout event of the monitoring timer and the network reply message event.

[0031] Create a dedicated event wait loop, and set timeout events and network reply message events to have the same processing priority in the dedicated event wait loop; within the dedicated event loop, wait for both timeout events triggered by the monitoring timer and network reply message events triggered by the network interface.

[0032] A dedicated event wait loop refers to a program logic loop that runs continuously at the system level and is specifically designed to handle specific types of events. Within this loop, the program suspends the current thread until it is awakened by one or more specific events registered in the loop. In this application, the dedicated event wait loop is specifically created to listen for two key events—timeout and network response—in parallel, serving as the core mechanism for implementing dual-channel monitoring.

[0033] A timeout event is a signal that is actively sent to the system by the monitoring timer after it reaches a preset initial threshold or target timeout period.

[0034] By configuring the system, a dedicated event waiting loop is created. Within this loop, timeout events and network reply message events are given the same processing priority. This means that the system no longer passively and solely waits for network replies, but instead waits in parallel within this dedicated event loop for either of these two events to occur first. This ensures that the system can process both network anomalies and normal replies with equal efficiency.

[0035] Step 103: In response to a timeout event being detected and no network reply message being detected, the first network driver is triggered to resend the network release message.

[0036] S1: In response to a timeout event being triggered within a dedicated event wait loop, check whether a valid network reply message has been received at the time the timeout event was triggered or within the dedicated event wait time.

[0037] S2: In response to not receiving a valid network reply message, determine that a timeout event has been detected and no network reply message event has been detected.

[0038] When a timeout event is triggered within the dedicated event waiting loop, the system first performs a status check to see if a valid network reply message has been successfully received within the just-ended waiting period, in order to eliminate unnecessary retransmissions caused by event response timing issues.

[0039] After confirming that no valid response has been received, the system determines that a timeout event has been detected but no network response message event has been detected.

[0040] S3: Increment the value of the retransmission counter. The retransmission counter is used to record the number of times the first network driver retransmits the network release message.

[0041] S4: Determine whether the current value of the retransmission counter is less than or equal to the preset retransmission threshold.

[0042] S5: In response to the current value of the retransmission counter being less than or equal to the preset retransmission threshold, a preset callback function is invoked, and the dedicated storage area is accessed through the preset callback function to obtain the protocol context of the network session persistently stored in the dedicated storage area.

[0043] Each time a network release message is retransmitted, the retransmission counter is incremented. This counter is dedicated to accurately recording the cumulative number of times the first network driver retransmits the network release message. The current value of the retransmission counter is compared with a preset retransmission threshold to determine if the retransmission attempt is still within the allowed limit. If the retransmission count has not exceeded the threshold, the system calls a preset callback function. This function accesses a dedicated storage area and retrieves the protocol context of the previously persistently saved network session, ensuring that retransmission is based on complete and consistent session information even in abnormal conditions.

[0044] S6: The preset callback function reconstructs and sends a target network release message with semantics consistent with the first network release message, based on the protocol context of the network session.

[0045] The pre-defined callback function uses the obtained protocol context to precisely reconstruct a target network release message that is semantically identical to the first message sent, and then resends it.

[0046] Specifically, the preset callback function reconstructs the network release message based on the protocol context of the network session; reads the protocol context of the first network session saved when the network release message was first sent from the dedicated storage area; determines whether the transaction identifier and device identifier in the reconstructed network release message match the protocol context of the first network session; and, in response to the matching of the transaction identifier and device identifier in the target network release message with the protocol context of the first network session, resends the reconstructed network release message as the target network release message.

[0047] The preset callback function is a function pre-registered by the system and automatically invoked after a specific event, such as a timeout event. In this application, it is configured to compare the core identifiers, such as the transaction identifier and device identifier, in the reconstructed network release message with the original identifiers stored in the initial network session protocol context before retransmitting the network release message. Only when these two key identifiers match completely is it proven that the retransmission is a legitimate continuation of the original session, and only then will the reconstructed message be sent as a valid target network release message. This ensures the correctness and security of the retransmission behavior and prevents protocol errors caused by context confusion.

[0048] The process of retransmitting the reconstructed network release message as the target network release message includes: calculating the target network release message transmission time based on network latency, network latency weight coefficient, link latency parameters, network link speed, estimated size of the target network release message, link transmission weight coefficient, server load coefficient, minimum timeout protection time, the last network release message transmission time, and the transmission time calculation formula; transmitting the target network release message based on the target network release message transmission time, and simultaneously updating the target timeout time of the monitoring timer corresponding to this target network release message transmission according to the target network release message transmission time.

[0049] Network latency in the network release message process refers to the average round-trip time from sending a network release message to receiving its acknowledgment reply message; it is used as a metric to measure network condition. It can be calculated by recording the message sending and receiving timestamps during protocol interactions.

[0050] The network latency weighting coefficient is a coefficient that is dynamically adjusted based on the stability and trend of network latency. It can be set by operators based on practical experience. Generally, the greater the network jitter, the larger the value of the network latency weighting coefficient should be, so as to reserve more buffer time.

[0051] Link delay parameter refers to the inherent propagation delay of a signal on a physical link, which is related to the physical distance and the medium.

[0052] It can be set by the operator based on practical experience, or it can be obtained by querying information such as LLDP of the network device. In the pre-boot environment, it can also be simplified to an experience value set according to the network type (such as LAN, WAN).

[0053] Network link speed refers to the connection rate of the network interface, which can be 100Mbps, 1Gbps, or 10Gbps. It can be obtained directly by querying the network card's hardware registers or by calling the network protocol interface provided by UEFI / BIOS.

[0054] The estimated size of the target network release message refers to the expected length of the DHCPv6 release message that will be retransmitted. Its size can be directly calculated and obtained by the program after it is constructed.

[0055] Link transmission weighting coefficients are weighting factors used to convert packet size and link speed into transmission time. They take into account protocol encapsulation overhead and hardware processing latency.

[0056] Server load factor can be a parameter used to assess the current processing capacity of a PXE / DHCPv6 server. Generally, the higher the load, the slower the expected response time may be. It can be inferred by analyzing trends in historical server response times; in more advanced implementations, the server can inform the client of its load status through specific options in DHCPv6 messages.

[0057] The minimum timeout protection time is a manually set minimum timeout threshold used to prevent the calculated transmission time from being too short and to avoid excessively frequent retransmissions. It can be set based on the system's minimum tolerable waiting interval or based on the operator's experience.

[0058] The last network release message sending time refers to the most recent time a network release message was sent. The system's timestamp can be recorded each time a message is sent and persistently stored in the network session context.

[0059] The target timeout is a new timeout period set after this retransmission to wait for a reply. It is usually associated with the time when the target network release message was sent and can be calculated based on the time when the target network release message was sent and the updated network delay.

[0060] The formula for calculating the transmission time is as follows: ; Among them, T n Indicates the target network release message sending time, RTT base K represents the network delay during the network release message process. rttThe network latency weighting coefficient is represented by L, the link latency parameter is represented by LinkSpeed, the network link speed is represented by S, and the estimated size of the target network release packet is represented by K. link T represents the link transmission weight coefficient, ServerLoad represents the server load coefficient, and T represents the link transmission weight coefficient. min T represents the minimum timeout protection time. n-1 This indicates the time when the last network release message was sent.

[0061] In this way, while ensuring that the minimum timeout protection time is not less than the minimum timeout protection time, the impact of network conditions (latency, bandwidth) and server real-time load on packet processing capacity is comprehensively considered, so that the retransmission interval dynamically adapts to the actual system load, which not only avoids network congestion, but also improves the efficiency and stability of the retransmission mechanism.

[0062] Step 104: In response to the number of times the network release message is retransmitted exceeds the preset retransmission threshold, the target event polling operation is forcibly terminated, and network control is transferred from the first network driver to the second network driver, so that the second network driver can take over and continue to execute the startup process of the pre-boot execution environment based on its own network protocol stack.

[0063] Specifically, the system retrieves the number of times the network release message has been retransmitted. If the number of times the network release message has been retransmitted exceeds a preset retransmission threshold, the system forcibly interrupts the target event polling operation of the first network driver. It then iterates through the network buffer pool held by the first network driver and releases the network buffer pool. The system queries the network protocol instance registered by the second network driver through the protocol interface. If the initialization state and operation state of the network protocol instance are both valid, the system determines that the network protocol interface of the second network driver is in a ready state and transfers network control from the first network driver to the second network driver.

[0064] The preset retransmission threshold is a pre-set maximum number of times that network release messages are allowed to be retransmitted, serving as a security mechanism to prevent the system from falling into an infinite retransmission loop.

[0065] A network buffer pool can be a storage area created by the first network driver specifically for storing network packets to be sent and received.

[0066] A network protocol instance refers to a specific, operable network protocol interface object registered with the system by the second network driver, representing the driver's readiness to provide network services.

[0067] The initialization status indicates that the hardware, memory, and other resources required for this network protocol instance have been successfully configured. The operation status indicates that the network protocol instance is active and can process network data streams normally.

[0068] Network control refers to the control over the configuration and management of the hardware resources (such as registers, DMA engine, and interrupts) of a network interface card.

[0069] When the number of retransmissions of network release messages exceeds a preset retransmission threshold, the system first forcibly terminates the target event polling operation of the first network driver to break a potential deadlock. Then, a thorough resource cleanup is performed, which involves traversing and releasing all network buffer pools held by the first network driver. After cleanup, the system queries the network protocol instances registered by the second network driver through the protocol interface and verifies their initialization and operational states. When it is confirmed that the instance is fully ready, the system determines that the network protocol interface of the second network driver is in a working state and then performs the final operation, formally transferring network control from the first network driver to the second network driver, enabling it to take over the network hardware and continue the pre-boot process.

[0070] This application also includes, before the forced termination of the target event polling operation: identifying the network session corresponding to the network release message that has been sent but has not received a valid network reply message corresponding to the first network driver; performing resource cleanup operations on the identified network session, including releasing the network buffer occupied by the identified network session and deregistering the monitoring timer corresponding to the identified network session.

[0071] This application also includes the following steps after the forced termination of the target event polling operation: generating a network session diagnostic report, which includes: the number of network release message transmissions, the number of invalid network reply messages received, the number of task priority conflicts, and network link quality indicators. The number of task priority conflicts is the cumulative number of network reply message losses caused by task priority competition between the first network driver and the second network driver when the monitoring timer timeout event is triggered; sending the diagnostic report to a remote log server or local persistent storage; and determining whether to trigger a network driver version rollback based on the diagnostic report.

[0072] In response to triggering a network driver version rollback, the system queries a pre-configured driver version repository for a historical version driver that is compatible with the current first network driver; it then uses the image file of the historical version driver to overwrite the image file of the current first network driver; it resets the configuration parameters of the first network driver to their initial default values; it reinitializes the first network driver and re-attempts the network release message sending process based on the reset configuration parameters.

[0073] Network link quality metrics are quantitative parameters used to assess the current network connectivity status, and may include link speed, signal strength, bit error rate, etc.

[0074] Network driver version rollback is a fault recovery strategy that refers to automatically replacing the current problematic first network driver with a known stable older version driver when a serious compatibility problem is diagnosed.

[0075] A driver version repository is a repository that stores multiple historical versions of network driver image files and their metadata (such as version number and compatibility information).

[0076] The initial default values ​​can be the original configuration parameters that the network driver has at the factory or during installation.

[0077] Before the forced termination operation, this application adds a precise resource reclamation mechanism. It identifies all incomplete network sessions initiated by the first network driver and performs thorough resource cleanup on these sessions, releasing their occupied network buffers and deregistering their dedicated monitoring timers, thus preventing resource leaks. After the forced termination operation, this application also adds an intelligent operation and maintenance diagnosis and recovery mechanism. It automatically generates a network session diagnostic report, which details key information such as the number of packet transmissions, the number of invalid replies, the inferred number of task priority conflicts, and network link quality indicators, and sends this report to remote or local storage. Based on this report, the system determines whether to trigger a network driver version rollback. If triggered, it queries a compatible historical driver version from the driver version repository, overwrites the current driver with its image file, resets the driver configuration to its initial default values, and reinitializes it. This automatically repairs startup failures caused by driver compatibility issues, greatly improving the system's self-healing capabilities and operation and maintenance efficiency.

[0078] In one embodiment, before transferring network control from the first network driver to the second network driver, the process includes: extracting digital signature data and a signer key identifier from a preset option field of the target network release message to be verified; querying a list of trusted public keys pre-stored in the firmware non-volatile storage area based on the signer key identifier to obtain a verification public key that matches the signer key identifier; performing a digital signature verification operation on the message payload portion of the target network release message using the verification public key, the message payload portion being the original data portion in the target network release message used to calculate the digital signature; in response to a successful digital signature verification operation, determining that the target network release message is complete and has a trusted source, and executing the step of transferring network control to the second network driver; in response to a failed digital signature verification operation, determining that the target network release message has a risk of tampering or an abnormal source, interrupting the transfer of network control, recording a security audit event, and triggering a security recovery process for the first network driver.

[0079] Digital signature data is ciphertext data appended to a message, generated by the sender using their private key to encrypt the message payload. It serves to prove the message's integrity and origin. It can be directly parsed and extracted from specific option fields (such as a custom DHCPv6 option) of the received target network release message.

[0080] The issuer key identifier is an identity identifier used to uniquely identify the private key used for signing, so that the recipient can quickly locate the correct verification public key from multiple public keys. It can be extracted from the option field of the message along with the digital signature data.

[0081] A trusted public key list is a pre-configured, trusted set of signing public keys and their corresponding key identifiers. Only messages signed with public keys from this list are trusted. This list can be pre-loaded into the firmware's non-volatile storage (such as UEFI variable storage) by the system administrator or device manufacturer and queried during verification.

[0082] The verification public key is the public key portion of the asymmetric key used to verify the validity of the digital signature. It can be obtained by matching and querying a list of trusted public keys based on the extracted issuer key identifier.

[0083] The message payload refers to the application layer data portion (i.e., the DHCPv6 message itself) remaining after removing headers that may change during transmission (such as IP headers and UDP headers) from the target network release message. This is the raw data used for signature calculation. It can be obtained by decoding and stripping the received raw network message according to the protocol format.

[0084] To ensure the security of control transfer, this application introduces a digital signature verification mechanism before the transfer. The system first extracts the digital signature data and the issuer's key identifier from the target network release message. Then, it queries a list of trusted public keys pre-stored in the firmware's secure storage area based on this identifier to obtain the corresponding verification public key. Next, the system uses this public key to perform digital signature verification on the message payload. Specifically, it uses the verification public key to perform verification calculations on the digital signature, and simultaneously performs hash calculations on the received message payload. Finally, it compares the calculation result with the hash value. If the comparison results match, the digital signature is deemed valid, the message is complete and its source is trustworthy, and the control transfer is allowed. If the comparison results do not match, verification fails, the message is considered to be at risk of tampering or forgery, the transfer process is immediately interrupted, and this anomaly is recorded as a security audit event. A security recovery process for the first network driver is initiated, effectively resisting man-in-the-middle attacks and malicious code injection, providing crucial security for the pre-boot environment.

[0085] In one feasible implementation, the preset retransmission threshold in this application can be an adaptive threshold dynamically adjusted based on historical diagnostic reports. The dynamic adjustment of the adaptive threshold includes: calculating an environmental stability coefficient based on the correlation between network link quality indicators and the number of task priority conflicts statistically analyzed in historical diagnostic reports; increasing the adaptive threshold to enhance the system's tolerance in harsh environments when the environmental stability coefficient is lower than a first threshold; and decreasing the adaptive threshold to optimize startup efficiency when the environmental stability coefficient is higher than a second threshold and the overall success rate of network release message transmission remains higher than a set level. Thus, the preset retransmission threshold is set as a self-optimizing adaptive threshold. By analyzing historical diagnostic reports, the correlation between network link quality and the number of task priority conflicts is calculated to obtain a quantified environmental stability coefficient. When this coefficient is low, it indicates that the system is in a complex network environment or experiencing frequent driver conflicts. In this case, the retransmission threshold is automatically increased to enhance the system's fault tolerance by allowing more retransmission attempts, ensuring a high startup success rate. Conversely, when the environment is stable and the success rate remains high, the retransmission threshold is decreased to avoid unnecessary waiting, thereby accelerating startup efficiency under normal conditions. This enables the system to intelligently perceive the environment and dynamically balance reliability and efficiency.

[0086] The pre-boot execution environment startup method provided in this application constructs a dual-channel event monitoring mechanism by introducing an independent monitoring timer and a dedicated event waiting loop, effectively solving the system deadlock problem caused by task priority competition between network drivers. Through intelligent retransmission and protocol context persistence, reliable recovery of the network release process is achieved in complex network environments; through digital signature verification and secure recovery processes, a secure protection system for the control transfer process is established; and through diagnostic reports and driver version rollback mechanisms, system self-diagnosis and self-repair capabilities are formed. These technical means work together to transform the traditionally fragile and vulnerable startup process into a robust system with fault tolerance, security, and self-healing capabilities, significantly improving the startup success rate and operational reliability in large-scale deployment scenarios.

[0087] In the server and cloud computing fields, PXE (Preboot Execution Environment), as a fundamental network boot technology, has become a core pillar of modern data center automated operation and maintenance. By shifting the server boot process from local storage to the network, it has fundamentally changed the traditional server deployment and maintenance model (the boot process of the Preboot Execution Environment). In hyperscale data centers, enterprise private clouds, and hybrid cloud environments, PXE technology enables administrators to deploy operating systems on hundreds of physical servers within minutes, improving efficiency by several orders of magnitude compared to traditional optical drive or USB flash drive installation methods.

[0088] With the rapid development of cloud computing and software-defined data centers, the role of PXE has expanded from simple operating system installation to the management of the entire server lifecycle. It not only supports common Linux and Windows system deployments but also enables automated configuration of cloud-native infrastructures such as Kubernetes nodes (an open-source container orchestration system) and OpenStack compute nodes (an open-source Infrastructure as a Service platform). In hyperconverged infrastructure (HCI), the integration of PXE with automation tools (such as Ansible and Terraform) makes unified allocation of compute, storage, and network resources possible. In enterprise IT operations and maintenance scenarios, the value of PXE is equally significant. By combining it with enhanced implementations such as iPXE (an open-source preboot execution environment), operations and maintenance personnel can build intelligent network boot menu systems to achieve diverse functions such as system repair, hardware diagnostics, and firmware upgrades. Especially in the event of server hardware failure, PXE can quickly load diagnostic tools, significantly shortening the mean time to repair (MTTR). In the security field, the combination of UEFI PXE (a preboot execution environment based on a unified extensible firmware interface) and Secure Boot ensures the security of the network boot process and prevents the injection of malicious code.

[0089] Compared to traditional IPv4 PXE (which uses the IPv4 protocol to boot PXE), IPv6 PXE (which uses the IPv6 protocol to boot PXE) demonstrates significant advantages in modern data center and cloud computing environments. Its 128-bit address space completely solves the address exhaustion problem, making it particularly suitable for the automated deployment needs of ultra-large-scale server clusters (the boot process after network control transfer). In terms of security, the natural integration of IPv6 PXE with IPsec (a protocol standard for providing encryption and security authentication for network communication over IP) provides end-to-end encryption, while perfectly compatible with the UEFI SecureBoot (Secure Boot based on a unified and extensible firmware interface) mechanism, building a complete trust chain from firmware to operating system. In hybrid cloud and edge computing scenarios, IPv6 PXE supports end-to-end network connectivity, enabling remote system installation and maintenance across geographical regions without the need for NAT (Network Address Translation) traversal. These technological advantages make IPv6-based PXE a key technological foundation for supporting the automated operation and maintenance of next-generation data centers, the mass deployment of IoT devices, and the construction of cloud-native infrastructure, providing new possibilities for the large-scale, intelligent, and secure operation of IT infrastructure.

[0090] During the iPXE network boot process in an IPv6 environment, a critical technical flaw exists when paired with UEFI BIOS (Basic Input / Output System) firmware: a probabilistic system crash occurs during the boot phase. The triggering mechanism for this issue is as follows: When the PXE client completes the NBP (Network Self-Test) file download and begins loading the iPXE program, iPXE initializes its own network driver and simultaneously uninstalls the original UEFI (Unified Extensible Firmware Interface) DHCPv6 driver (the first network driver's uninstallation process is triggered by the loading of the second network driver). At this time, the UEFI DHCPv6 (Unified Extensible Firmware Interface Dynamic Host Configuration Protocol version 6) driver proactively sends a DHCPv6 RELEASE message to the PXE server (the PXE client proactively releases the IPv6 address assigned by the PXE server, transmitted via UDPv6 protocol, ports 546 / 547). This message contains the IPv6 address to be released and the server identifier (UUID). Under normal circumstances, the PXE server should reply with a DHCPv6 REPLY message (the PXE server replies with a DHCPv6 REPLY message to confirm the release of the IPv6 address, also using the UDPv6 protocol, as it is suitable for connectionless scenarios with high real-time requirements), confirming the release of the IPv6 address and updating the address pool status.

[0091] In actual operation, the UEFI firmware polls network packets by raising the TPL (Task Priority Level) to the TPL_CALLBACK level (callback-level network task priority level), while the iPXE driver also sets its own TPL to the same priority, causing resource contention between the two (the implementation mechanism of the target event polling operation). That is, during the UEFI IPv6 PXE startup process, when the iPXE takes over network control, the UEFI native DHCPv6 driver, after sending a DHCPv6 RELEASE message, enters a polling state to wait for the server's DHCPv6 REPLY (the event polling operation where the first network driver waits for the network reply message corresponding to the network release message). Because the iPXE driver raises its task priority (TPL) to the same TPL_CALLBACK level as the UEFI driver, resource contention occurs, and critical DHCPv6 REPLY messages may be dropped. The UEFI driver, unable to receive a response, remains blocked in the polling state, ultimately leading to a system startup deadlock.

[0092] In the PXE boot process of the UEFI architecture, after the client server completes the hardware initialization (including memory, CPU, and basic chipset configuration) of the BIOS PEI (Pre-EFInitialization) stage, it loads and executes most of the hardware drivers in the DXE (Driver Execution Environment) stage, and the system enters the BDS (Boot Device Selection) stage. At this time, the BIOS enumerates all network interfaces in sequence and identifies the network ports that support PXE boot by detecting the PCI Option ROM or the UEFI PXE-BC protocol (a standard protocol built into the UEFI firmware for implementing network boot (PXE) functionality). For each candidate network port, the BIOS loads the native UEFI DHCPv6 driver, establishes an IPv6 link-local address (FE80:: / 10), and obtains the global IPv6 address, boot server information (OPT_BOOTFILE_SERVER - boot file server option), and network boot program path (OPT_BOOTFILE_URL - boot file URL option) through complete DHCPv6 interactions (Solicit, Advertise, Request, Reply). After successfully downloading the NBP (Network Boot Program) file and transferring control to iPXE, iPXE initializes its high-performance network driver (based on UEFISNP - Simple Network Protocol) and notifies the BIOS to uninstall the native UEFI DHCPv6 driver via ExitBootServices(). At this point, the unloaded UEFI DHCPv6 driver will proactively send a DHCPv6 RELEASE message (using UDPv6 port 547) to release the temporarily allocated IPv6 address and enter Polling mode to wait for the server's DHCPv6 REPLY confirmation (also via UDPv6 port 546). Because iPXE raises the task priority level (TPL) to TPL_CALLBACK when taking over the network stack, a priority conflict arises with the UEFI driver's event polling mechanism, causing the critical DHCPv6 REPLY message to be discarded. This TPL contention keeps the UEFI driver blocked in the Polling state, unable to complete the resource release process, ultimately causing a deadlock in the system startup process.The core of the problem lies in the incompatibility between the UEFI network stack and the third-party network driver (iPXE) in terms of task priority management and resource release timing.

[0093] To address this issue, this application encapsulates the process of the BIOS unloading the native UEFI DHCPv6 driver (including the process of sending DHCPv6 RELEASE messages and receiving DHCPv6 REPLY messages in polling mode) into a reentrant callback function, which is then integrated into the DHCPv6Stop function (starting a monitoring timer independent of the target event polling operation). Specifically, after the PXE client successfully sends the DHCPv6 RELEASE message, the system immediately starts a configurable timer (tested on a server with this issue; an initial recommended setting is 5 seconds). This timer runs in parallel with the existing polling status detection mechanism (monitoring the timer's startup and monitoring mechanism). When the timer triggers a timeout event, it indicates that the PXE client may be experiencing polling state blockage due to TPL (task priority) contention between the BIOS and iPXE. The system will automatically call back the DHCPv6Stop function to re-initiate the complete DHCPv6 RELEASE process (including message construction, sending, and status maintenance) (triggering the first network driver to resend the network release message). To improve reliability, we designed an intelligent retransmission strategy: a multiplicative backoff algorithm is used to dynamically adjust the retransmission interval (e.g., 5 seconds for the first time, 10 seconds for the second, and 15 seconds for the third), and the consistency of the transaction ID and server identifier (DUID - device unique identifier) ​​is strictly verified during each retransmission. Simultaneously, the system maintains a retransmission counter. When the preset maximum number of retransmissions (3 recommended) is reached, even if no DHCPv6 REPLY message is received, it will safely assume that the PXE server has completed the IPv6 address release. At this point, the system will: forcibly exit the polling loop, clean up the residual state of the UEFI network protocol stack, and completely transfer network control to the iPXE network driver (forcibly terminating the target event polling operation and transferring network control to the second network driver).

[0094] This application considers boundary case handling, including: verifying the iPXE driver readiness state before final forced exit, saving necessary DHCPv6 session information for fault diagnosis, and ensuring the correct release of memory resources. This mechanism not only solves the deadlock problem during startup caused by TPL contention but also enhances robustness in complex network environments. It ensures the normal startup of the PXE environment for numerous PXE clients in customer server cluster scenarios, providing more reliable underlying support for UEFI IPv6 PXE network startup and significantly improving the IPv6 PXE startup success rate. The intelligent retransmission strategy, while ensuring the completion of IPv6 address release, does not significantly increase PXE startup time (around 30 seconds), which is within an acceptable range. This improves system robustness in complex PXE network environments and solves the compatibility issue between UEFI network drivers and iPXE network drivers.

[0095] Please see Figure 2 In a specific application scenario, when the first network driver is a UEFI native DHCPv6 driver and the second network driver is an iPXE driver, in order to solve the compatibility problem between the UEFI network driver and the iPXE network driver, the startup method of the pre-boot execution environment of this application can be as follows: 1. During the PXE boot process in a UEFI architecture, after the client server completes the key hardware initialization of the BIOS PEI (Pre-EFInitialization) phase (including memory controller configuration, CPU microcode update, basic chipset register settings, and PCIe device enumeration), the system officially enters the DXE (Driver Execution Environment) phase. In this phase, the BIOS performs deep initialization on each detected network interface controller by traversing the PCIe device tree, including: 1) Read the Option ROM information through the PCI configuration space to verify PXE compatibility (compliant with PXE-BC2.1 or later specifications).

[0096] The system accesses a specific register area within the network card's PCI configuration space, reads its preset Option ROM (Option Read-Only Memory) code, parses the firmware identifier and instruction set in the ROM, and rigorously verifies whether it complies with Intel's PXE-BC 2.1 or higher technical specifications, thereby confirming that the network card hardware has basic PXE boot capabilities.

[0097] 2) For network cards that conform to the UEFI specification, load the device-specific UEFI driver (following the UNDI (Universal Network Driver Interface) protocol).

[0098] For UEFI-compliant network cards that have passed compatibility verification, the system locates and loads a device driver tailored for that network card model from its Option ROM or firmware built-in driver library. This driver is developed in strict accordance with the UNDI (Universal Network Driver Interface) protocol, and its loading provides a standardized and unified low-level network hardware operation interface for upper-layer software, enabling the initialization and basic control of the network card.

[0099] 3) Establish a complete UEFI network protocol stack (including basic protocols such as ARPv6 (Address Resolution Protocol for IPv6) and ICMPv6 (Internet Control Message Protocol for IPv6)).

[0100] After the network card driver is successfully loaded and the hardware is initialized, the UEFI firmware will build a simplified network protocol stack running in the pre-boot environment on top of its driver layer. This protocol stack fully implements core network protocols such as ARPv6 (for IPv6 address resolution) and ICMPv6 (for network diagnostics and control), laying the necessary protocol foundation for subsequent DHCPv6 communication, address acquisition, and network boot program download.

[0101] 2. After successfully recognizing the PXE boot port, the system will instantiate the UEFI native DHCPv6 driver. This driver first establishes an IPv6 link-local address (FE80:: / 10) through RS (Router Solicitation) / RA (Router Advertisement) interaction, and then initiates a complete DHCPv6 four-way handshake process: 1) Send a Solicit message (FF02::1:2) to probe for available servers.

[0102] After obtaining the link-local address, the PXE client broadcasts a Solicit message via IPv6 multicast (destination address FF02::1:2, representing all DHCPv6 servers and relay agents). This message is equivalent to a network probe, designed to discover available DHCPv6 servers on the network and request basic network configuration and startup parameters.

[0103] 2) Parse Option 59 (Boot Server Identifier) ​​in the Advertise message.

[0104] Upon receiving a Solicit message, one or more DHCPv6 servers in the network will respond with an Advertise message. The client parses these response messages, and a key operation is extracting the Option 59 (Startup Server Identifier) ​​option. This option specifies the address of a particular PXE server that can provide the client with the startup files, allowing the client to select a server for subsequent requests.

[0105] 3) Request specific configurations via Request messages.

[0106] Based on the selection made in the previous step, the client sends a Request message to the specified PXE server (as determined by Option 59). This message formally requests specific network configurations from the server, including the global IPv6 address, DNS information, and crucial parameters such as the location of the boot file (OPT_BOOTFILE_URL).

[0107] 4) Finally, obtain the Reply message from the IPv6 PXE server to complete the PXE service registration, which will facilitate the subsequent loading of the iPXE driver.

[0108] After processing the Request message, the selected PXE server sends a final Reply message. The client receives this message and obtains its assigned global IPv6 address, the exact boot file path, and other configuration information. This signifies the completion of the DHCPv6 interaction and successful PXE service registration. The client can then use the obtained URL address to download and execute network boot programs such as iPXE via TFTP or HTTP.

[0109] 3. During the NBP (Network Bootloader) loading phase (such as iPXE.efi), the system undergoes a complex control transfer process: 1) The BIOS loads the NBP image using LoadImage().

[0110] The UEFI BIOS uses the LoadImage() service to load the network bootloader from the storage device or network into system memory. This step completes the loading of the NBP image and basic memory and format verification, preparing it for execution, but the NBP code has not yet started running.

[0111] 2) Call StartImage() (start image) to transfer execution permissions.

[0112] After successfully loading the image, the BIOS calls the StartImage() service. This call formally transfers CPU execution control from the UEFI firmware to the NBP's entry function, marking the start of independent operation for network boot programs such as iPXE.

[0113] 3) During iPXE initialization, it registers its own SNP (Simple Network Protocol) implementation through InstallProtocolInterface().

[0114] During initialization, iPXE calls the InstallProtocolInterface() function. This operation registers its own implementation of the SNP instance with the UEFI system, thereby declaring its readiness to take over network functions and override or supplement the original Simple Network Protocol of UEFI.

[0115] 4) Call ExitBootServices() (exit startup service) to trigger UEFI environment cleanup.

[0116] Once iPXE confirms that it is fully ready and capable of independently managing all necessary hardware and resources, it calls the ExitBootServices() function. This call is a critical turning point in the system's operational state, notifying the UEFI firmware that the operating system bootloader (such as iPXE) is ready. UEFI can then release most of the runtime service resources it occupies, thereby triggering the unloading process of UEFI native drivers (including the DHCPv6 driver).

[0117] 4. At this point, the iPXE network driver will completely take over hardware control, including: 1) Reset the network card's DMA (Direct Memory Access) engine.

[0118] The iPXE driver first resets the network card's DMA engine. This step aims to stop all ongoing data transfers by the network card, clear the DMA cache queue, and ensure that the data channel between the network card and the host memory is in a new, defined initial state, laying the foundation for exclusive control of iPXE.

[0119] 2) Reconfigure the receive descriptor ring.

[0120] iPXE releases or reallocates the receive descriptor ring set by the UEFI driver and establishes a new descriptor ring structure according to its own memory management strategy and performance requirements. This allows the network card to directly store received network packets into the memory area specified by iPXE according to the instructions of the iPXE driver, thereby completely taking over the network data reception process.

[0121] 3) Update MAC (Media Access Control) filter settings.

[0122] iPXE reconfigures the MAC address filtering register of the network card according to the needs of its network protocol stack. This includes setting new unicast addresses, updating multicast group lists, etc., to ensure that the network card only receives network packets destined for the local machine or those of interest to iPXE, while filtering out irrelevant data, optimizing performance, and ensuring the correctness of communication. 4) Guides the UEFI to execute the native DHCPv6 driver exit mechanism.

[0123] After the iPXE driver completes full takeover and initialization of the network hardware, it guides the UEFI to execute the exit mechanism of its native DHCPv6 driver. This is the starting point for triggering the address release process, marking the imminent relinquishment of the UEFI network stack and the formal handover of final network control responsibilities to the ready iPXE driver.

[0124] 5. When UEFI executes the native DHCPv6 driver exit mechanism, it strictly follows the RFC 8415 specification: 1) Construct a RELEASE message containing (Identity Association Identifier) ​​and DUID (Device Unique Identifier).

[0125] The driver first assembles a DHCPv6 RELEASE message conforming to RFC 8415 in memory. The message must contain the critical identification fields: IAID (Identity Association Identifier), which specifies the particular address lease to be released; and DUID (Device Unique Identifier), used to uniquely identify the client device to the server.

[0126] 2) Set the transaction ID (xid) to match the previous session.

[0127] Obtain the unique transaction ID for this session from the context of the previously successful DHCPv6 session. Set the obtained transaction ID into the corresponding field of the newly constructed RELEASE message. This step is crucial to ensure that the PXE server can correctly associate and process this release request with the session that initially assigned the address.

[0128] 3) Send to All_DHCP_Relay_Agents_and_Servers (FF02::1:2) via UDPv6 port 547, wait for the PXE server to return a DHCPv6 REPLY message, and update the UDPv6 status information of the BIOS.

[0129] The assembled RELEASE message is sent from the client's port 547 to the target address FF02::1:2 (the link-local multicast address for all DHCPv6 relay agents and servers) via the UDPv6 protocol. After the message is sent, the driver enters a waiting state, expecting to receive a DHCPv6 REPLY acknowledgment message from the server via port 546. Once a valid reply is received, the driver updates the UDPv6 protocol status information within the BIOS to reflect that the address has been successfully released.

[0130] 6. During the waiting period for a REPLY, the system faces multi-level resource contention. This contention can prevent the UEFI from successfully obtaining DHCPv6 REPLY messages through the network card and updating the BIOS's UDPv6 status information, leading to a deadlock during the IPv6 PXE boot process. At the network card event handling level: the ISR (Interrupt Service Routine) registered by the iPXE driver will preempt the UEFI polling loop; since both TPLs are set to TPL_CALLBACK (0x08), data packets are lost.

[0131] While the UEFI firmware continuously checks the network interface card (NIC) for data arrival via polling, the loaded iPXE driver registers its interrupt service routine (ISR) with the system. When a DHCPv6 REPLY message arrives at the NIC and triggers an interrupt, the iPXE ISR will preemptively gain execution rights because it has the same task priority as the UEFI polling loop. This can cause packets that should be processed by the UEFI stack to be incorrectly intercepted or discarded by the iPXE stack, resulting in the UEFI polling loop never receiving the expected packet.

[0132] At the memory management level: UEFI uses network buffers allocated by BootServices; iPXE may use its own memory pool for management; buffer ownership conflicts can lead to packet dropping.

[0133] The UEFI network stack uses its BootServices to allocate and manage memory buffers for storing network packets. However, the iPXE driver, during initialization and hardware takeover, establishes its own independent memory pool to manage these network buffers. When iPXE begins resetting the network interface card and reconfiguring the DMA engine, the ownership of the buffer held by UEFI becomes ambiguous; it may have been reclaimed or marked invalid by iPXE. Therefore, even if a packet successfully arrives in host memory, it may become invalid because it has been written to an "unclaimed" or overwritten buffer, preventing UEFI from reading it correctly.

[0134] To address the issue of the UEFI DHCPv6 driver failing to receive DHCPv6 REPLY messages due to resource contention during the iPXE driver loading-triggered unloading process, a parallel monitoring mechanism is established by starting a monitoring timer independent of the UEFI event polling loop after the DHCPv6 RELEASE message is sent. This mechanism monitors both timer timeout events and network reply message events. When a timeout occurs and no reply message is received, the system automatically triggers the DHCPv6 release message retransmission process, with the number of retransmissions controlled by a retransmission counter. After reaching the preset maximum number of retransmissions, the system forcibly terminates the waiting process, completing the transfer of network control to the iPXE driver. By setting a dynamically reentrant callback function architecture, the UEFI DHCPv6 driver unloading process is reconstructed, establishing a dual-channel retransmission and exit mechanism. This architecture transforms the traditional linear execution process into an intelligent retransmission system with state preservation capabilities, achieving reliable recovery capabilities during the unloading process while maintaining protocol specifications. First, the independent monitoring channel and intelligent retransmission mechanism effectively solve the system deadlock problem caused by task priority contention, significantly improving the success rate of IPv6 PXE startup. Second, the dual-channel monitoring architecture ensures that the system can recover from deadlock under any abnormal conditions, greatly enhancing the system's robustness. Finally, the forced exit mechanism provides the ultimate guarantee for the transfer of control, ensuring that the startup process can be reliably completed under any network conditions.

[0135] Please see Figure 3 , Figure 3 A schematic diagram of the PXE environment connection is shown, such as... Figure 3 As shown, the system consists of PXE clients (servers to be deployed), a PXE server, and network infrastructure. Multiple PXE clients are connected to the PXE server via a network switch, forming a local area network (LAN).

[0136] PXE servers typically integrate DHCPv6 and TFTP / HTTP boot services (file services). When a PXE client powers on, its network card broadcasts a request to the network. This request is captured by the PXE server, which then assigns an IPv6 address to the client via DHCPv6 and informs it of the location of the boot file (NBP). Finally, the client downloads and executes the boot file via TFTP or HTTP, thus initiating the operating system installation or boot process. In other words, during PXE boot, the client obtains the necessary resources and configurations from a remote server over the network, rather than from its local disk.

[0137] Please see Figure 4 , Figure 4 The PXE server architecture is shown, such as Figure 4As shown, the PXE server architecture includes a PXE server, which carries DHCPv6 service (IPv6 Dynamic Host Configuration Protocol service), file storage service, and network interface layer.

[0138] A PXE server can be deployed on a physical machine or a virtual machine. Its core consists of two working service components: a DHCPv6 service and a file service. The DHCPv6 service handles dynamic address allocation requests from clients, while the file service (TFTP / HTTP Boot) stores and transmits files such as the Network Bootloader (NBP) and operating system images to clients. The DHCPv6 service and file storage service can be deployed on the same physical server or associated through configuration (e.g., the DHCPv6 service specifies the file server's address and boot file path in its response).

[0139] The DHCPv6 service listens for DHCPv6 requests (Solicit / Request) from clients and responds with a message containing the IPv6 address, next-hop server address (Option 59), and boot file URL (Option 60). The file service receives file transfer requests initiated by clients based on the information in the DHCPv6 response and sends the necessary NBP (such as iPXE.efi) or kernel image to the client. The PXE server architecture is the context in which the problem occurs; its DHCPv6 service needs to correctly receive and process DHCPv6 RELEASE messages from the client's UEFI stack and respond with REPLY messages.

[0140] Please see Figure 5 , Figure 5 The PXE client architecture is shown in Figure 5. Under the UEFI BIOS environment, the PXE service client performs PXE booting via IPv6 network in four logical stages: hardware initialization, network configuration, file download, and system takeover.

[0141] Hardware initialization includes: Firmware initialization: After the computer is powered on, the UEFI firmware first performs a self-test and hardware initialization. Establishing the UEFI BIOS environment: Loading UEFI drivers and services to provide the operating environment for subsequent operations. Loading the network card ROM: Executing the PXE firmware code on the network card to enable network boot capabilities. Initializing the SNP network protocol stack: Loading a simple network protocol for the UEFI environment, which is the foundation for network communication. Loading the PCIe UNDI driver: Standardizing the control of the network card hardware through a generic network device interface driver. Building the UEFI network protocol stack: Finally, establishing a complete UEFI network service, preparing for the next stage of obtaining network configuration.

[0142] Network configuration includes: UEFI DHCPv6 client startup: The DHCPv6 client running in the UEFI environment begins operation. Sending SOLICIT messages: The client broadcasts a DHCPv6 request to the network, searching for an available server. Processing ADVERTISE / REPLY messages: Receives and processes the advertising and acknowledgment messages returned by the DHCPv6 server. Communicating via UDPv6 ports 546-547: Completes the entire communication interaction with the DHCPv6 server. Parsing the OPTION_BOOTFILE option: Parses the URL address of the next startup file (NBP) to be downloaded from the server's reply.

[0143] The file download includes: Downloading the NBP file: Based on the URL obtained in the previous stage, a network bootloader, such as iPXE or GRUB2, is downloaded via TFTPv6 or HTTPv6 protocol. Loading iPXE / GRUB2: A more powerful bootloader is loaded and run in the client's memory to support more complex operations and protocols. Obtaining the kernel and initrd: The bootloader continues to download the operating system's kernel image and initial memory disk via TFTPv6 or HTTPv6 protocol according to the configuration. Communication using IPv6 GUA / ULA addresses: Throughout the file download process, the client communicates with the server using the IPv6 global unicast address or unique local address obtained in the second stage.

[0144] System takeover includes: Driver loading and iPXE driver takeover: Loading a higher-performance custom network driver (such as the iPXE driver) to prepare for taking over network control. Unloading native UEFI drivers: Proactively disabling and uninstalling native network drivers that come with UEFI and have poor performance or compatibility. Reconfiguring the network card DMA engine: Reconfiguring the network card's direct memory access engine to optimize data transmission efficiency and prepare for the operating system to take over the hardware. Calling the ExitBootServices function: Officially informing the UEFI firmware that the boot service has ended, after which the operating system kernel will fully take over control of the hardware. SNP protocol reloading: In the new driver environment, reinitializing the network protocol stack to ensure the continuity and stability of network functions before the operating system boots.

[0145] The PXE boot process consists of four core phases: First, in the address acquisition phase, the client obtains network configuration and boot parameters from the server via a DHCPv6 four-way handshake (Solicit-Advertise-Request-Reply); next, in the program loading phase, the client downloads the network boot program via TFTP / HTTP protocol based on the obtained information; then, in the critical control transfer phase, after the iPXE program is loaded and ExitBootServices() is called, the system enters a deadlock while waiting for confirmation of the address release request due to task priority contention; finally, in the boot phase, after successfully resolving the deadlock, the iPXE driver reconfigures the network and loads the operating system image to complete the boot process.

[0146] This application systematically refactors the DHCPv6Stop function, transforming the traditional linear execution flow into a fault-tolerant intelligent retransmission mechanism. This implementation addresses the reliability issue of IPv6 address release in a UEFI environment through a layered design. Its main technical improvements include the following key aspects: 1. Modular and State Machine Design: The previously rigid sequential execution flow is decoupled into an independent callback function using a state machine pattern. This function clearly defines five core states: initialization, sending, listening, retransmission, and termination. Specifically: Initialization: Preparatory work, setting a 5-second initial timer and resetting the retransmission counter to zero. Send: Constructing and sending a DHCPv6 release message conforming to RFC8415. Listen: Waiting for a server response via the UEFI network protocol stack's event notification mechanism. Retransmission: Handling timeout events and executing intelligent retransmission strategies. Termination: Completing resource cleanup and final transfer of control. This makes the entire release process a robust, traceable, and recoverable process.

[0147] 2. Intelligent Timing and Retransmission Strategy: An intelligent timer system is introduced to dynamically manage timeout waiting. The strategy is to set the initial wait time to 5 seconds, and if a timeout occurs, the interval between subsequent retransmissions will increase progressively. This gradual rollback approach effectively avoids increasing the burden on the network during periods of congestion, while providing more time for system recovery.

[0148] 3. Dual Safeguards and Control Logic: To ensure absolute security, the system maintains a dual safeguard mechanism: a retransmission counter strictly limits the maximum number of retransmissions to 3 to prevent infinite loops. Furthermore, before each retransmission, the consistency of critical information such as the transaction ID and server identifier is rigorously verified to ensure that each retransmission is valid and correct. Clear termination conditions are also established. Once a response is successfully received, the retransmission limit is reached, or an iPXE driver anomaly is detected, the system will immediately terminate the waiting, forcibly exit the polling loop, clean up the protocol stack state, and transfer control to iPXE, thereby completely breaking the deadlock.

[0149] In this way, by modularizing the release process into an intelligent state machine, and supplementing it with an adaptive retransmission strategy and a robust protection mechanism, the release of IPv6 PXE addresses is transformed from a fragile link into a highly reliable, automatically fault-tolerant, and robust process.

[0150] This application features a deep optimization and restructuring of the execution process, resulting in a three-stage, logically rigorous workflow: In the initialization phase, the system first establishes the basic environment, including registering key callback functions, allocating dedicated storage for persistent protocol states, and starting the first round of monitoring timers, laying the foundation for subsequent fault-tolerant processing. Upon entering the core polling phase, the process first sends a DHCPv6 RELEASE message, then enters an event-driven waiting loop. In this loop, the system monitors both network responses and timers in parallel. When a timeout event occurs, the system does not become blocked but activates intelligent processing logic to dynamically decide whether to initiate a retransmission or execute subsequent strategies, ensuring the process always progresses. Finally, in the termination phase, the system performs comprehensive cleanup operations: thoroughly clearing the network buffer to prevent resource leaks, rigorously verifying that network control has been completely transferred to the iPXE driver, and writing critical debugging information to the logs. This closed-loop design ensures that regardless of the path the process takes, the system reaches a stable and predictable final state.

[0151] The embodiments of this application provide a startup device for a pre-boot execution environment, the startup device for the pre-boot execution environment specifically as follows: Figure 6 As shown, the startup device for the pre-startup execution environment includes: startup module 20, monitoring module 21, sending module 22, and execution module 23.

[0152] The startup module 20 is used to respond to the first network driver triggering the target unloading process due to the loading of the second network driver, and the first network driver has sent a network release message in the target unloading process, to start a monitoring timer independent of the target event polling operation. The target event polling operation is the event polling operation executed by the first network driver waiting for the network reply message corresponding to the network release message.

[0153] The monitoring module 21 is used to respond to the start of the monitoring timer and to monitor the timeout event of the monitoring timer and the network reply message event.

[0154] The sending module 22 is used to trigger the first network driver to resend the network release message in response to a timeout event being detected and no network reply message being detected.

[0155] The execution module 23 is used to forcibly terminate the target event polling operation in response to the number of times the network release message is retransmitted exceeds the preset retransmission threshold, and transfer network control from the first network driver to the second network driver, so that the second network driver can take over and continue to execute the startup process of the pre-boot execution environment based on its own network protocol stack.

[0156] like Figure 7 As shown, embodiments of this application also provide a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above-described pre-boot execution environment startup method embodiments.

[0157] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described pre-boot execution environment startup method embodiments at runtime.

[0158] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0159] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0160] The above provides a detailed description of the startup method for a pre-boot execution environment provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only intended to help understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A method for starting a pre-boot execution environment, characterized in that, The startup method of the pre-boot execution environment includes: In response to the first network driver triggering the target unloading process due to the loading of the second network driver, and the first network driver having sent a network release message in the target unloading process, a monitoring timer independent of the target event polling operation is started. The target event polling operation is the event polling operation executed by the first network driver waiting for the network reply message corresponding to the network release message. In response to the start of the monitoring timer, the system also monitors the timeout event of the monitoring timer and the network reply message event. In response to the detection of the timeout event and the absence of the network reply message event, the first network driver is triggered to resend the network release message; In response to the number of times the network release message is retransmitted exceeds a preset retransmission threshold, the target event polling operation is forcibly terminated, and network control is transferred from the first network driver to the second network driver, so that the second network driver can take over and continue to execute the startup process of the pre-boot execution environment based on its own network protocol stack.

2. The startup method for the pre-boot execution environment according to claim 1, characterized in that, The step of starting a monitoring timer independent of the target event polling operation in response to the first network driver triggering the target unloading process due to the loading of the second network driver, and the first network driver having sent a network release message during the target unloading process, includes: During the initialization of the target unloading process, a dedicated storage area is configured to persistently save the protocol context of the current network session. The network session is the entire process of the first network driver sending a network release message and receiving a network reply message corresponding to the network release message. The protocol context includes key parameters used to uniquely identify this session, the client device, and the network resources to be released. The key parameters include a transaction identifier and a device unique identifier. Call the firmware interface to create a monitoring timer based on a hardware clock source; In response to the completion of the network release message transmission, the monitoring timer is initialized and started, and the timeout period of the monitoring timer is set to the initial threshold. The interrupt response of the monitoring timer is not affected by the task priority level of the first network driver.

3. The startup method for the pre-boot execution environment according to claim 1, characterized in that, The response to the start of the monitoring timer, and the monitoring of the timeout event and network reply message event of the monitoring timer, includes: Create a dedicated event wait loop, and set the timeout event and the network reply message event to have the same processing priority in the dedicated event wait loop; Within the dedicated event loop, it simultaneously waits for a timeout event triggered by the monitoring timer and a network reply message event triggered by the network interface.

4. The startup method for the pre-boot execution environment according to claim 1, characterized in that, The step of triggering the first network driver to resend the network release message in response to the detection of the timeout event and the absence of the network reply message event includes: In response to the timeout event being triggered within the dedicated event wait loop, check whether a valid network reply message has been received at the timeout event was triggered or within the dedicated event wait time. In response to the failure to receive a valid network reply message, it is determined that the timeout event has been monitored and the network reply message event has not been monitored. The value of the retransmission counter is incremented, and the retransmission counter is used to record the number of times the first network driver retransmits the network release message; Determine whether the current value of the retransmission counter is less than or equal to the preset retransmission threshold; In response to the current value of the retransmission counter being less than or equal to the preset retransmission threshold, a preset callback function is invoked, and the dedicated storage area is accessed through the preset callback function to obtain the protocol context of the network session persistently stored in the dedicated storage area. The preset callback function reconstructs and sends a target network release message that is semantically consistent with the first network release message sent, based on the protocol context of the network session.

5. The startup method for the pre-boot execution environment according to claim 4, characterized in that, The preset callback function, based on the protocol context of the network session, reconstructs and sends a target network release message semantically consistent with the initially sent network release message, including: The preset callback function reconstructs the network release message based on the protocol context of the network session; Read the initial network session protocol context saved when the network release message was first sent from the dedicated storage area; Determine whether the transaction identifier and device identifier in the reconstructed network release message match the initial network session protocol context; In response to the transaction identifier and device identifier in the target network release message matching the initial network session protocol context, the reconstructed network release message is retransmitted as the target network release message.

6. The startup method for the pre-boot execution environment according to claim 5, characterized in that, The step of retransmitting the reconstructed network release message as the target network release message includes: The target network release message sending time is calculated based on network latency, network latency weight coefficient, link latency parameters, network link speed, estimated size of the target network release message, link transmission weight coefficient, server load coefficient, minimum timeout protection time, last network release message sending time, and the sending time calculation formula. The target network release message is sent based on the target network release message sending time, and the target timeout time of the monitoring timer corresponding to this target network release message sending is updated according to the target network release message sending time.

7. The startup method for the pre-boot execution environment according to claim 6, characterized in that, The formula for calculating the transmission time is expressed as follows: ; Among them, T n Indicates the target network release message sending time, RTT base K represents the network delay during the network release message process. rtt The network latency weighting coefficient is represented by L, the link latency parameter is represented by LinkSpeed, the network link speed is represented by S, and the estimated size of the target network release packet is represented by K. link T represents the link transmission weight coefficient, ServerLoad represents the server load coefficient, and T represents the link transmission weight coefficient. min T represents the minimum timeout protection time. n-1 This indicates the time when the last network release message was sent.

8. The startup method for the pre-boot execution environment according to claim 1, characterized in that, The step of forcibly terminating the target event polling operation and transferring network control from the first network driver to the second network driver in response to an operation that retransmits a network release message more than a preset retransmission threshold includes: The number of times the network release message is retransmitted is obtained. In response to the number of times the network release message is retransmitted exceeding a preset retransmission threshold, the target event polling operation driven by the first network is forcibly interrupted. Iterate through the network buffer pool held by the first network driver and release the network buffer pool; Query the network protocol instances registered by the second network driver through the protocol interface installation service; In response to the fact that both the initialization state and the operation state of the network protocol instance are valid, it is determined that the network protocol interface of the second network driver is in a ready state, and network control is transferred from the first network driver to the second network driver.

9. The startup method for the pre-boot execution environment according to claim 1, characterized in that, Before forcibly terminating the target event polling operation, the following steps are also included: Identify the network session corresponding to the network release message that has been sent but has not received a valid network reply message, which corresponds to the first network driver; Perform resource cleanup operations on the identified network sessions, including releasing the network buffer occupied by the identified network sessions and deregistering the monitoring timers corresponding to the identified network sessions.

10. The startup method for the pre-boot execution environment according to claim 1, characterized in that, The forced termination of the target event polling operation includes: Generate a network session diagnostic report, which includes: the number of network release messages sent, the number of invalid network reply messages received, the number of task priority conflicts, and network link quality indicators. The number of task priority conflicts is the cumulative number of network reply messages lost due to task priority competition between the first network driver and the second network driver when the timeout event of the monitoring timer is triggered. The diagnostic report is sent to a remote log server or local persistent storage; Based on the diagnostic report, determine whether to trigger a network driver version rollback.

11. The startup method for the pre-boot execution environment according to claim 10, characterized in that, The determination of whether to trigger a network driver version rollback includes: In response to triggering a network driver version rollback, the system queries a pre-configured driver version repository for a historical version driver that is compatible with the current primary network driver. Use the image file driven by the previous version to overwrite the image file of the current first network driver; Reset the configuration parameters of the first network driver to their initial default values; The first network driver is reinitialized, and the network release message sending process is retried based on the reset configuration parameters.

12. The startup method for the pre-boot execution environment according to claim 1, characterized in that, The process of transferring network control from the first network driver to the second network driver includes: Extract the digital signature data and the issuer key identifier from the preset option field of the target network release message to be verified; Based on the issuer key identifier, query the list of trusted public keys pre-stored in the firmware non-volatile storage area to obtain the verification public key that matches the issuer key identifier; Using the verification public key, a digital signature verification operation is performed on the payload portion of the target network release message, wherein the payload portion is the original data portion in the target network release message used to calculate the digital signature; In response to the successful digital signature verification operation, it is determined that the target network release message is complete and has a trustworthy source, and the step of transferring network control to the second network driver is executed; In response to the failure of the digital signature verification operation, it is determined that the target network release message is at risk of tampering or has an abnormal source, and the transfer process of network control is interrupted, a security audit event is recorded, and a security recovery process for the first network driver is triggered.

13. A computer device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the startup method of the pre-boot execution environment as described in any one of claims 1 to 12 when executing the computer program.

Citation Information

Patent Citations

  • Method and device for sending message in Ethernet dual-homed link protection

    CN101741535A

  • Network and server resource management method and system thereof

    CN106255145A

  • Methods and apparatus for session release in wireless communication

    CN110582995A

  • Network resource release method, network element, communication system and storage medium

    CN118139217A

  • Server cooperative control method, storage medium and electronic equipment

    CN119917350A

Cited By

  • Computing equipment starting method, computing equipment starting device, medium and program product

    CN122308941A

  • Computer device startup methods, computer device startup devices, media and program products

    CN122308941B