Network information detection method and device, chip, network interface card, equipment, medium and program product
By recording the sending and receiving time in the probe request message and calculating the network delay, the problem of delaying the load of the HOST CPU is solved, and the accuracy and efficiency of network delay calculation are improved.
Patent Information
- Application Number
- CN202510617105.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art is unable to process CQE in time when calculating network delays, resulting in increased network delays and RDMA network failures are difficult to locate and troubleshoot.
By generating a detection request message carrying the detection start time, the dispatch network card records the sending and receiving time during transmission, and records these times in the response message to calculate the accurate network delay.
It improves the accuracy of network delay calculation, reduces CPU consumption, and all detection information is stored in messages, and there is no need to save detection-related information locally or remotely.
Smart Images

Figure CN120238470A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communication technologies, and in particular, to a network information detection method, apparatus, chip, network interface card, device, medium, and program product. Background Art
[0002] As the number of parameters of AI large models increases, the demand scale for the network in the intelligent computing center becomes larger, and the throughput also becomes larger. The traditional TCP network protocol stack relies on the processing of the CPU and is difficult to meet the requirements of AI training in terms of latency and throughput performance. RDMA (Remote Direct Memory Access) offloads the protocol processing flow to the hardware and supports the hardware to directly access service data, reducing the data transfer latency. There is a great improvement in the overall throughput and latency performance. However, the RDMA network also faces the problem that it is difficult to locate and troubleshoot network failures in the data center network. And the TCP service network fault detection tool Pingmesh cannot detect the fault problems of the RDMA network because the traditional TCP runs in a lossy network, while RDMA uses a lossless network. In order to reduce the direct mutual interference between TCP and RDMA traffic, it is usually necessary to isolate TCP and RDMA by running them in different Traffic Class queues.
[0003] However, the current method relies on the HOST to record the timestamp after the RDMA network card reports CQE, and then calculates the network latency. Once the HOST CPU load is relatively heavy and it cannot process CQE in time, it will cause an additional increase in network latency. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a network information detection method, apparatus, chip, network interface card, device, medium, and program product that can improve the accuracy of network latency calculation.
[0005] In a first aspect, the present application provides a network information detection method, which is applied to a first server. The method includes:
[0006] Generate a detection request message, where the detection request message carries a detection start time;
[0007] Schedule a first network card to transmit the detection request message to a second network card, and during the transmission of the detection request message, the detection request message sending time and the detection request message receiving time are recorded in the detection request message;
[0008] Receive a probe response message corresponding to the probe request message. The probe response message is transmitted by the second network card to the first server via the first network card, and during the transmission process of the probe response message, the network information determined based on the probe start time, the probe request message sending time, and the probe request message receiving time recorded in the probe request message is recorded in the probe response message.
[0009] In one embodiment, the scheduling of the first network card to transmit the probe request message to the second network card and recording the probe request message sending time and the probe request message receiving time in the probe request message during the transmission process of the probe request message includes:
[0010] Schedule the first network card to send a probe request message to the second network card and fill in the probe request message sending time in the probe request message; the probe request message is used to instruct the second network card to fill in the probe message receiving time in the probe request message after receiving the probe request message.
[0011] In one embodiment, the receiving of the probe response message corresponding to the probe request message includes:
[0012] Receive the probe response message through the first network card, where the probe response message received by the first network card carries the network card queuing processing time of the second network card obtained based on the probe request message receiving time, and the first network card obtains the network transmission delay based on the probe response message receiving time, the network card queuing processing time, and the probe request message sending time, and fills the network transmission delay into the probe response message;
[0013] Based on the time when the first server receives the probe response message, the probe start time, the network transmission delay, and the network card queuing processing time, obtain the first server processing time.
[0014] In one embodiment, the generating of the probe request message includes:
[0015] Based on the list of servers to be probed information, determine the second server information;
[0016] Generate a probe request message based on the second server information and record the probe start time.
[0017] In one embodiment, before the scheduling of the first network card to transmit the probe request message to the second network card, it further includes:
[0018] Based on the second server information of the second server to be probed, determine the queue and end-to-end link of the extended reliable datagram type;
[0019] The dispatching of the first network card to transmit the detection request message to the second network card includes:
[0020] Sending doorbell information to the first network card, where the doorbell information is used to instruct the first network card to transmit the detection request message to the second network card based on the determined queue of the extended reliable datagram type and the end-to-end link.
[0021] In one embodiment, the network information includes at least one of the network card queuing processing time of the second network card, the network transmission delay, and the first server processing time; the method further includes:
[0022] Determining network anomalies based on at least one of the following:
[0023] Obtaining the congestion situation of the second network card based on the change situation of the network card queuing processing time of the second network card;
[0024] Determining the congestion situation of the network switch between the first network card and the second network card based on the network transmission delay; and
[0025] Determining the load situation of the first server based on the first server processing time.
[0026] In a second aspect, the present application further provides a network information detection method, which is applied to the second network card, and the method includes:
[0027] Receiving a detection request message sent by the first network card, where the detection request message is generated by the first server, and the detection request message carries a detection start time, and the detection request message sending time and the detection request message receiving time are recorded in the detection request message during the transmission process of the detection request message;
[0028] Generating a detection response message corresponding to the detection request message, and transmitting the detection response message to the first server via the first network card. During the transmission process of the detection response message, the network information determined based on the detection start time, the detection request message sending time, and the detection request message receiving time recorded in the detection request message is recorded in the detection response message.
[0029] In one embodiment, the generating a detection response message corresponding to the detection request message, and transmitting the detection response message to the first server via the first network card includes:
[0030] Construct an empty probe response message, copy the probe start time and the probe request message sending time in the probe request message to the empty probe response message, obtain the network card queuing processing time of the second network card based on the probe request message reception time, and fill the network card queuing processing time into the probe response message;
[0031] Transmit the probe response message to the first server via the first network card. When the first network card receives the probe response message, obtain the network transmission delay based on the probe response message reception time, the network card queuing processing time, and the probe request message sending time, and fill the network transmission delay into the probe response message. Then transmit the probe response message filled with the network transmission delay to the first server. The first server is used to obtain the first server processing time based on the time when the first server receives the probe response message, the probe start time, the network transmission delay, and the network card queuing processing time.
[0032] In one embodiment, the probe request message includes the probe request message sending time and the probe request message reception time. The probe request message sending time is filled in the probe request message when the first network card sends the probe request message, and the probe request message reception time is filled in the probe request message when the second network card receives the probe request message.
[0033] In one embodiment, receiving the probe request message sent by the first network card includes:
[0034] Receiving the probe request message sent by the first network card based on the determined queue of the extended reliable datagram type and the end-to-end link; the queue of the extended reliable datagram type and the end-to-end link are determined by the first network card based on the second server information of the second server to be probed.
[0035] In one embodiment, the network information includes at least one of the network card queuing processing time of the second network card, the network transmission delay, and the first server processing time; the change situation of the network card queuing processing time of the second network card is used to characterize the congestion situation of the second network card; the network transmission delay is used to determine the congestion situation of the network switch between the first network card and the second network card; the first server processing time is used to determine the load situation of the first server.
[0036] In a third aspect, the present application further provides a network information detection device, and the device includes:
[0037] A detection request message generation module, configured to generate a detection request message, where the detection request message carries a detection start time;
[0038] A first message sending module, configured to schedule a first network card to transmit the detection request message to a second network card, and during the transmission of the detection request message, the detection request message sending time and the detection request message receiving time are recorded into the detection request message;
[0039] A first message receiving module, configured to receive a detection response message corresponding to the detection request message, where the detection response message is transmitted by the second network card to the first server via the first network card, and during the transmission of the detection response message, network information determined based on the detection start time, the detection request message sending time, and the detection request message receiving time recorded in the detection request message is recorded into the detection response message.
[0040] In a fourth aspect, the present application further provides a network information detection device, where the device includes:
[0041] A second message receiving module, configured to receive a detection request message sent by a first network card, where the detection request message is generated by a first server, and the detection request message carries a detection start time, and during the transmission of the detection request message, the detection request message sending time and the detection request message receiving time are recorded into the detection request message;
[0042] A second message sending module, configured to generate a detection response message corresponding to the detection request message, and transmit the detection response message to the first server via the first network card, and during the transmission of the detection response message, network information determined based on the detection start time, the detection request message sending time, and the detection request message receiving time recorded in the detection request message is recorded into the detection response message.
[0043] In a fifth aspect, the present application further provides a chip, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the method in any one of the above embodiments are implemented.
[0044] In a sixth aspect, the present application further provides a network interface card, including the chip in any one of the above embodiments and multiple interfaces, and the chip processes data or communicates externally through the interfaces.
[0045] In a seventh aspect, the present application further provides a computer device, including the network interface card in any one of the above embodiments, and the network interface card is used to process data or communicate externally.
[0046] In an eighth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method in any one of the above embodiments are implemented.
[0047] In a ninth aspect, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method in any one of the above embodiments are implemented.
[0048] For the above network information detection method, device, chip, network interface card, device, medium and program product, a detection request message is generated, and the detection request message carries a detection start time; the first network card is scheduled to transmit the detection request message to the second network card, and during the transmission process of the detection request message, the detection request message sending time and the detection request message receiving time are recorded in the detection request message; a detection response message corresponding to the detection request message is received, the detection response message is transmitted by the second network card to the first server via the first network card, and during the transmission process of the detection response message, the network information determined based on the detection start time, the detection request message sending time and the detection request message receiving time recorded in the detection request message is recorded in the detection response message. In this way, the first network card, the second network card and the review server all participate in the recording and calculation of the detection, the detection granularity is finer, the detection accuracy is higher, the detection result is more accurate, and all detection information is stored in the message, and no detection-related information needs to be saved locally and remotely. When the first server receives the detection response message, it does not need to query the status of the detection request, and the corresponding remote second server does not need to participate in the detection process, reducing the consumption of the CPU during the detection process. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can be obtained based on these drawings.
[0050] Figure 1 It is an architecture diagram of Pingmesh in the traditional technology;
[0051] Figure 2 It is a schematic diagram of the RC method in the traditional technology;
[0052] Figure 3 It is a schematic diagram of the UD method in the traditional technology;
[0053] Figure 4Schematic diagram of the network detection process corresponding to the RC method in the traditional technology;
[0054] Figure 5 Schematic diagram of the network detection process corresponding to the UD method in the traditional technology;
[0055] Figure 6 Schematic diagram of the working framework of the network information detection method in one embodiment;
[0056] Figure 7 Schematic flow chart of the network information detection method in one embodiment;
[0057] Figure 8 Timing diagram of the network information detection method in one embodiment;
[0058] Figure 9 Schematic diagram of the request scheduling model in one embodiment;
[0059] Figure 10 Schematic flow chart of the network information detection method in another embodiment;
[0060] Figure 11 Structural block diagram of the network information detection device in one embodiment;
[0061] Figure 12 Structural block diagram of the network information detection device in another embodiment;
[0062] Figure 13 Internal structure diagram of a computer device in one embodiment. Detailed implementation manners
[0063] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0064] With the rapid development of cloud computing and AI large models, the scale of data center networks (DCNs) is becoming larger and larger. Various large-scale distributed servers are built on top of physical data centers, such as distributed file systems, distributed storage systems, and distributed computing systems. Many components of these systems need to interact through the network within the data center or across data centers. In such large-scale systems, software and hardware failures are the norm rather than the exception. Locating and recovering data center network failures faces several challenges.
[0065] Challenge 1: It is difficult to determine that it is a "network" problem after a failure: Some components in the distributed system are intermittently accessed, or the end-to-end delay suddenly increases at the 99th percentile, and the network throughput suddenly drops. Generally, these will be classified as "network" failures. However, in practice, about 50% of these "network" failures are not really caused by the network. It may be caused by reasons such as CPU overload and end-side hardware failures.
[0066] Challenge 2: Define and track network service level agreements (SLAs). Many services require the network to provide certain performance guarantees. For example, in AI training services, the training speed depends on the last packet response of the slowest-performing server. These services are very sensitive to network latency and packet loss. When there are a large number of services in the network, it becomes very difficult to provide separate network measurement and tracking for different services.
[0067] Challenge 3: Network troubleshooting. When a network failure occurs, it will affect the interests of customers and partners and needs to be detected, mitigated, and resolved as soon as possible. However, a data center network consists of hundreds of thousands to millions of servers, hundreds of thousands and millions of cables and optical fibers. Therefore, it is a difficult task to determine where the problem lies.
[0068] To quickly identify the network fault location problem in the data center, Pingmesh was proposed in traditional technologies for data center network delay measurement and analysis of fault causes. Combined Figure 1 as shown Figure 1 is the architecture diagram of Pingmesh in traditional technologies. Pingmesh includes three modules: a control module (Pingmesh Controller), an agent module (Pingmesh Agent), and a data storage and analysis module (Data storage and Analysis). The Pingmesh agent module needs to be deployed on all servers. Its main task is to download the server ping list from the controller, conduct network probes on each server in the list one by one, and then upload the results to the data storage and analysis module. To distinguish and perceive the reasons for the increased delay, ICMP or UDP cannot be used for ping probes, but TCP or HTTP closer to the business is used for probing. The Pingmesh control module is responsible for generating the Ping list for all servers and running an algorithm to decide which servers should ping which other servers. The data storage and analysis module is mainly responsible for collecting and saving the Ping results and comprehensively analyzing the network fault points and fault causes based on the Ping results of all servers.
[0069] As the number of parameters in AI large models increases, the demand scale for the network in intelligent computing centers becomes larger and the throughput also becomes larger. The traditional TCP network protocol stack relies on CPU processing and is difficult to meet the requirements of AI training in terms of latency and throughput performance. Since the protocol processing flow of RDMA (Remote Direct Memory Access) is offloaded to hardware and it supports direct hardware access to service data, the data transfer latency is reduced. There is a significant improvement in the overall throughput and latency performance. However, the RDMA network also faces the problem of difficult fault location and troubleshooting in the data center network. Moreover, the TCP service network fault detection tool Pingmesh cannot detect RDMA network faults because traditional TCP runs in a lossy network while RDMA uses a lossless network. To reduce the direct mutual interference between TCP and RDMA traffic, TCP and RDMA usually need to be run in different Traffic Class queues for isolation.
[0070] The current mainstream data transfer methods of RDMA are two methods: RC (Reliable Connect) and UD (Unreliable Datagram). The RC method can be referred to Figure 2 as shown, and the UD method can be referred to Figure 3 as shown. Among them, RC is a reliable data transfer based on a connection. Both parties exchange data based on a unique pair of QPs (Queue Pair). RC supports both send (depending on the remote CPU processing) and write (not depending on the remote CPU processing) requests. UD is a connectionless and unreliable data transfer method. The local queue QP can send messages to any remote QP or receive messages sent by any remote QP.
[0071] To use the Pingmesh architecture to achieve a full - range monitoring of the RDMA network quality, either use the RC transmission mode: The monitoring server confirms the number of servers N in the data center and the number of network transmission paths M between the monitoring server and other servers, and creates N*M - 1 RC connections. Then use different RC connections to initiate RDMA network probes to other servers. The specific network probing process can be combined with Figure 4 as shown, including: 1) The monitoring server Host initiates a probe request prob and records the start time T0; 2) The RDMA network card of the monitoring server sends a write message with a length of 0; 3) After the RDMA network card of the remote server receives the write message, it replies with an ACK packet; 4) After the RDMA network card of the monitoring server receives the ACK, it reports CQE. After the HOST receives the CQE, it records the timestamp T1 and calculates the network latency RTT = T1–T0 to complete the network probe.
[0072] Either use the UD transmission mode: create a QP of the UD type, and use the same QP to perform RDMA network probing for all other services. The specific network probing process can be combined with Figure 5 as shown in the figure, including: 1) The monitoring server Host initiates a probing request and records the time T0; 2) The RDMA network card of the monitoring server sends a send message and reports the CQE at the same time. After the HOST receives the CQE, it records the time T1; 3) After the RDMA network card of the remote server receives the send message, it reports the CQE. After the remote server HOST receives the CQE, it records the time T2 and replies with a network card probing response; 4) The remote RDMA network card sends a send message; 5) After the RDMA network card of the monitoring server receives the send message, it reports the CQE; 6) After the monitoring server HOST receives the CQE, it records the timestamp T4, calculates the RTT = T4 – T0, and calculates the network delay RTT’ = T4 - T1 – (T3 – T2) to complete the probing process.
[0073] However, the above method has at least the following problems:
[0074] 1) The network probing is inaccurate: It depends on the HOST to record the timestamp after the RDMA network card reports the CQE and then calculate the network delay. Once the HOST CPU load is relatively heavy and cannot process the CQE in time, it will cause an additional increase in the network delay.
[0075] 2) Excessive resource occupation leads to performance degradation: Using the RC mode for RDMA network probing requires consuming a large amount of link resources: for N nodes and M paths, N*M – 1 connections need to be created. The more the number of connections, the lower the overall performance of RDMA.
[0076] 3) It cannot truly reflect the end-to-end delay of the service: The real service uses the RC traffic model, and it will be affected by flow control and congestion control when the RDMA network card sends data, but UD is not affected by flow control and congestion. This results in some deviation between the measurement results in the UD mode and the real network delay of the service.
[0077] 4) It cannot distinguish the cause of the failure: UD is unreliable data. When the local RDMA network card is overloaded and under excessive pressure, it may abandon sending UD data packets. When the remote RDMA network card has an abnormality or the HOST server is not ready to receive data, it may silently discard the UD data packets. At this time, the monitoring server cannot perceive it at all and will misjudge it as a network problem.
[0078] To solve one or more of the above problems, a network information probing method is proposed in this application, which can be applied to such as Figure 6The working framework shown. The working framework includes a first server, a first network card, a second network card, and a second server. The first network card corresponds to the first server, and the second network card corresponds to the second server. Each server in the data center can act as the first server, responsible for detecting the network status of all paths from the first server to all other second servers. The network detection pool of the first server stores the destination server addresses and path information (Destination IP And Path ID, DIPPID) to be detected. The detection request module polls all the server addresses and path information DIPPID in the pool, initiates a network detection request, and records the moment T0. The scheduling algorithm module calculates which QP (Queue Pair) and EEC (End-To-END Connection) need to be used for detecting the detection message according to the DIPPID information, and issues a DB (doorbell) to the first network card (monitoring RDMA network card). The first network card sends a detection request message to the network after receiving the DB and records the time T1. After receiving the network detection request message, the remote message receiving module of the second network card records the reception time T2, queues the message, and sends it to the detection response module for processing. The detection response module assembles the detection response message. The remote message sending module inserts the remote RDMA network card queuing processing time (Response ProcessTime, RSP_TIME) into the detection response message and then sends the response message. The monitoring RDMA network card records the time T3 after receiving the detection response message. The network delay calculation module calculates the network transmission delay (Network Round-Trip Time, NET_RTT) and reports it to the first server. The host delay calculation module obtains the network card reported information, records the current time as T4, and calculates the host delay (HostProcess Time, HOST_TIME). The network evaluation module evaluates the cause of the "network" failure according to the segment delay during the detection process: 1) The change in the remote RDMA network card delay reflects the congestion situation of the remote network card. An increase in the processing delay indicates an increase in the number of messages received by the remote network card. When it increases to a certain limit, packet loss may occur. 2) The network transmission delay reflects the congestion situation of the network switch. An increase in the transmission delay indicates network congestion. When it increases to a certain extent, packet loss may occur. 3) The host delay reflects the CPU load of the host.
[0079] In an exemplary embodiment, as Figure 7 shown, a network information detection method is provided. Taking the first server in Figure 1 as an example, the following steps 702 to 706 are included. Among them:
[0080] S702: Generate a detection request message, and the detection request message carries the detection start time.
[0081] The detection request message is generated by the first server, which is the server that performs network detection. Each server in the data center can act as the first server to detect the network status of all paths from the first server to all other second servers.
[0082] Specifically, the detection request message can include at least four fields. Optionally, each field can be four bytes. Specifically, the four fields are the detection request header REQ, the detection start time, the detection request message sending time, and the detection request message receiving time. The detection start time is the time obtained by the first server when it initiates the detection request and filled into the detection request message. The detection request message sending time is the time when the first network card corresponding to the first server sends the detection request message, and the detection request message receiving time is the time when the second network card corresponding to the second server receives the detection request message.
[0083] In this application, the detection request message is initiated by the first server, and the first server fills in the detection request header and the detection start time.
[0084] S704: Schedule the first network card to transmit the detection request message to the second network card. During the transmission process of the detection request message, the detection request message sending time and the detection request message receiving time are recorded into the detection request message.
[0085] Combined with Figure 8 as shown, where Figure 8 is a timing diagram of the network information detection method in an embodiment. The first server schedules the first network card to transmit the detection request message to the second network card. When the first network card receives the doorbell signal, the first network card obtains the detection request message sending time, fills it into the detection request message, and sends the detection request message. It should be noted that the difference between the detection request message at this time and the detection request message generated by the first server is that the detection request message sending time in the detection request message sent by the first network card is no longer empty.
[0086] After the second network card receives the detection request message, it obtains the detection request message receiving time and records the detection request message receiving time into the detection request message. It should be noted that the difference between the detection request message at this time and the detection request message sent by the first network card is that the detection request message receiving time in the detection request message processed by the second network card is no longer empty.
[0087] Therefore, during the transmission process of the detection request message, the time information of the detection request message during the transmission process is recorded.
[0088] In other embodiments, since there is at least one switch between the first network card and the second network card, when the probe request message arrives at each switch, the time when the switch receives the probe request message can also be recorded, and when the switch forwards the probe request message, the time when the switch sends the probe request message can be recorded. Thus, when it is determined that the network anomaly is caused by a network switch anomaly, the location of the specific switch corresponding to the network anomaly can be determined based on the data in the message.
[0089] S706: Receive a probe response message corresponding to the probe request message. The probe response message is transmitted by the second network card to the first server via the first network card, and during the transmission of the probe response message, the network information determined based on the probe start time, the probe request message sending time, and the probe request message receiving time recorded for the probe request message is recorded in the probe response message.
[0090] After receiving the probe request message, the second network card generates a corresponding probe response message. The probe response message corresponds to the probe request message and at least includes four fields. Optionally, each field includes four bytes. Specifically, the four fields include a probe response header, a first position corresponding to the probe start time, a second position corresponding to the probe request message sending time, and a third position corresponding to the probe request message receiving time.
[0091] During the transmission of the probe response message, it needs to pass through the second network card, the first network, and the first server. When passing through each entity, the entity generates corresponding network information based on the information carried in the probe response message, so that each entity in the network participates in the detection of network information.
[0092] Specifically, the probe response message is initiated by the second network card. The second network card copies the probe start time of the probe request message to the first position, copies the probe request message sending time in the probe request message to the second position, determines the network information corresponding to the third position based on the probe request message receiving time in the probe request message, and stores it in the third position.
[0093] Subsequently, the second network card sends the probe response message to the first network card. The first network card updates the network information at the second position based on the probe request message sending time, the time when it receives the probe response message, and the network information at the third position in the probe response message to obtain a new probe response message. The difference between the probe response message at this time and the probe response message sent by the second network card lies in the different network information corresponding to the second position.
[0094] Subsequently, the first network card sends the detection response message with the updated network information at the second location to the first server, and the first server updates the network information at the first location based on the detection start time, the time when the detection response message with the updated network information at the second location is received, and the network information at the second and third locations.
[0095] In this way, all the detection information is stored in the message, and no detection-related information needs to be saved locally and remotely. When the first server receives the detection response message, it does not need to query the status of the detection request. Moreover, compared with other solutions that use RDMA send semantics for detection, this application uses RDMA write semantics for end-to-end network detection, reducing the impact on the second server.
[0096] For the above network information detection method, a detection request message is generated, and the detection request message carries the detection start time; the first network card is scheduled to transmit the detection request message to the second network card, and during the transmission of the detection request message, the detection request message sending time and the detection request message receiving time are recorded in the detection request message; the detection response message corresponding to the detection request message is received. The detection response message is transmitted by the second network card to the first server via the first network card, and during the transmission of the detection response message, the network information determined based on the detection start time, the detection request message sending time, and the detection request message receiving time recorded in the detection request message is recorded in the detection response message. In this way, the first network card, the second network card, and the server all participate in the recording and calculation of the detection, the detection granularity is finer, the detection accuracy is higher, the detection result is more accurate, and all the detection information is stored in the message. No detection-related information needs to be saved locally and remotely. When the first server receives the detection response message, it does not need to query the status of the detection request, and the corresponding second server at the remote end does not need to participate in the detection process, reducing the consumption of the CPU during the detection process.
[0097] In one optional embodiment, scheduling the first network card to transmit the detection request message to the second network card and recording the detection request message sending time and the detection request message receiving time in the detection request message during the transmission of the detection request message includes: scheduling the first network card to send the detection request message to the second network card and filling in the detection request message sending time in the detection request message; the detection request message is used to instruct the second network card to fill in the detection message receiving time in the detection request message after receiving the detection request message.
[0098] Among them, in combination with Figure 8As shown, the first server schedules the first network card to transmit a detection request message to the second network card. When the first network card receives a doorbell signal, the first network card obtains the transmission time of the detection request message, fills it into the detection request message, and then sends the detection request message. It should be noted that the difference between the detection request message at this time and the detection request message generated by the first server is that the transmission time of the detection request message in the detection request message sent by the first network card is no longer empty.
[0099] After the second network card receives the detection request message, it obtains the reception time of the detection request message and records the reception time of the detection request message in the detection request message. It should be noted that the difference between the detection request message at this time and the detection request message sent by the first network card is that the reception time of the detection request message in the detection request message processed by the second network card is no longer empty.
[0100] In this way, during the transmission process of the detection request message, each participating entity participates in the recording and calculation of the detection, with finer detection granularity and higher detection accuracy.
[0101] In one alternative embodiment, receiving a detection response message corresponding to the detection request message includes: receiving the detection response message through the first network card, where the detection response message received by the first network card carries the network card queuing processing time of the second network card obtained based on the reception time of the detection request message, and the first network card obtains the network transmission delay based on the reception time of the detection response message, the network card queuing processing time, and the transmission time of the detection request message, and fills the network transmission delay into the detection response message; obtaining the first server processing time based on the time when the first server receives the detection response message, the detection start time, the network transmission delay, and the network card queuing processing time.
[0102] The transmission process of the detection response message is from the second network card to the first network card, and then from the first network card to the first server, and corresponding network information is recorded during the transmission process.
[0103] In one alternative embodiment, the network information includes at least one of the network card queuing processing time of the second network card, the network transmission delay, and the first server processing time; the method further includes: determining network anomalies based on at least one of the following: obtaining the congestion situation of the second network card based on the change situation of the network card queuing processing time of the second network card; determining the congestion situation of the network switch between the first network card and the second network card based on the network transmission delay; and determining the load situation of the first server based on the first server processing time.
[0104] The probe response message is initiated by the second network card. The second network card copies the probe start time of the probe request message to the first position, copies the probe request message sending time in the probe request message to the second position, determines the network information corresponding to the third position based on the probe request message receiving time in the probe request message, and stores it in the third position. The second network card obtains the network card queuing processing time based on the probe request message receiving time in the probe request message and the time for generating the probe response message, that is, the network information at the third position. The network card queuing processing time rsp_time = timestamp1 – T2, where timestamp1 is the time for generating the probe response message and T2 is the probe request message receiving time. The network card queuing processing time can reflect the congestion condition of the second network card. When this processing time becomes larger, it indicates that the number of received messages by the second network card increases, and packet loss may occur when it increases to a certain limit.
[0105] Subsequently, the second network card sends the probe response message to the first network card. The first network card updates the network information at the second position based on the probe request message sending time, the time for receiving the probe response message, and the network information at the third position in the probe response message to obtain a new probe response message. The difference between the probe response message at this time and the probe response message sent by the second network card lies in the different network information corresponding to the second position.
[0106] The first network card updates the network information at the second position, that is, the network transmission delay, based on the probe request message sending time, the time for receiving the probe response message, and the network information at the third position in the probe response message. The network transmission delay RTT = timestamp2 – T1 – rsp_timep, where timestamp2 is the time for the first network card to receive the probe response message and T1 is the probe request message sending time. The network transmission delay is used to reflect the congestion condition of the switch in the network. When the network transmission delay increases, it indicates that the switch in the network has a congestion condition.
[0107] Optionally, since there is at least one switch between the first network card and the second network card, when the probe response message reaches each switch, the time for the switch to receive the probe response message can also be recorded, and when the switch forwards the probe response message, the time for the switch to send the probe response message can be recorded. Thus, when it is determined that the network anomaly is caused by a network switch anomaly, the position of the specific switch corresponding to the network anomaly can be determined based on the data in the message. It should be noted that in this embodiment, the probe response message further includes the processing time of the corresponding switch in the probe request message.
[0108] Subsequently, the first network card sends the detection response message with the updated network information at the second location to the first server. The first server updates the network information at the first location based on the detection start time, the time when it receives the detection response message with the updated network information at the second location, and the network information at the second and third locations.
[0109] Among them, the first server updates the network information at the first location based on the detection start time, the time when it receives the detection response message with the updated network information at the second location, and the network information at the second and third locations. That is, the first server processing time. The first server processing time host_rtt = timestamp3 – T0 – RTT – rsp_time, where timestamp3 is the time when the first server receives the detection response message, and T0 is the detection start time. The first server processing time is used to reflect the CPU load situation of the first server.
[0110] In the above embodiments, more clearly and completely distinguishing the processing times of service messages in each module helps to quickly and accurately locate the cause of the "network failure". Moreover, in this application, both the detection request message and the detection response message only require an additional 14 bytes, reducing the impact of detection messages on the network. Both the first network card and the second network card participate in the recording and calculation process of detection, with a finer detection granularity, higher detection accuracy, and more accurate detection results.
[0111] In one optional embodiment, generating a detection request message includes: determining second server information based on the list of servers to be detected; generating a detection request message based on the second server information and recording the detection start time.
[0112] The first server includes a network detection pool, which stores the second server addresses and path information (Destination IP And Path ID, DIPPID) to be detected. The detection request module polls and traverses all the second server addresses and path information DIPPID in the network detection pool, initiates a network detection request, and records the moment T0.
[0113] In one optional embodiment, before scheduling the first network card to transmit the detection request message to the second network card, it further includes: determining the queue of the extended reliable datagram type and the end-to-end link based on the second server information of the second server to be detected; scheduling the first network card to transmit the detection request message to the second network card includes: sending a doorbell message to the first network card, and the doorbell message is used to instruct the first network card to transmit the detection request message to the second network card based on the determined queue of the extended reliable datagram type and the end-to-end link.
[0114] Among them, in combination with Figure 9 as shownFigure 9 Schematic diagram of a request scheduling model in an embodiment, where data is transmitted between servers through QP (Queue Pair) and EEC (End-To-End Connection). One EEC link is required for communication between every two servers, and one QP is required for each path of the server message outlet. EEC is a reliable transmission-based link, which has the same priority as the RC link used by the service for scheduling queues and policies. When sending messages, it is restricted by congestion control and flow control, and will initiate message retransmission after detecting packet loss. QP is an end-to-end link queue, which is divided into three types: RC (Reliable Connection), UD (Unreliable Datagram), and XRD (Extend Reliable Datagram). In this application, using the XRD type of QP can not only ensure end-to-end reliability but also achieve many-to-many communication between QPs (one QP can send messages to any other QP using EEC and can also receive messages sent by any other QP). When a probe request needs to be initiated, the probe algorithm module first selects an XRD QP (QueuePare) according to the path and selects EEC according to the second server IP address.
[0115] Then the first server issues a DB (doorbell) to the first network card. After receiving the DB, the first network card sends a probe request message to the network and records the sending time of the probe request message.
[0116] In the above embodiment, probing is performed based on the reliable QP. After the probe request message is lost, it will automatically initiate retransmission to ensure that valid probe results can be obtained for each probe without the intervention of the HOST. And the request scheduling model using QP plus EEC requires only N connections and M QPs for N remote servers and M paths, while using ordinary RC probing requires N*M–1 links. The more links there are, the higher the cost and the lower the performance. Therefore, this application improves the processing performance. In addition, the XRD type of QP and the end-to-end reliable link EEC are used to probe the RDMA service path. The same EEC link is used for probing between every two servers, and the same QP is used for probing the same outlet path.
[0117] In an exemplary embodiment, as Figure 10 shown, a network information probing method is provided. Taking the method applied to the second network card in Figure 6 as an example, it includes the following steps 1002 to step 1004. Among them:
[0118] S1002: Receive the probe request message sent by the first network card. The probe request message is generated by the first server and carries the probe start time. During the transmission of the probe request message, the probe request message sending time and the probe request message receiving time are recorded in the probe request message.
[0119] S1004: Generate a probe response message corresponding to the probe request message and transmit the probe response message to the first server via the first network card. During the transmission of the probe response message, the network information determined based on the probe start time, the probe request message sending time, and the probe request message receiving time recorded in the probe request message is recorded in the probe response message.
[0120] In one optional embodiment, generating a probe response message corresponding to the probe request message and transmitting the probe response message to the first server via the first network card includes: constructing an empty probe response message, copying the probe start time and the probe request message sending time in the probe request message to the empty probe response message, obtaining the network card queuing processing time of the second network card based on the probe request message receiving time, and filling the network card queuing processing time into the probe response message; transmitting the probe response message to the first server via the first network card. When the first network card receives the probe response message, obtain the network transmission delay based on the probe response message receiving time, the network card queuing processing time, and the probe request message sending time, and fill the network transmission delay into the probe response message. Transmit the probe response message filled with the network transmission delay to the first server. The first server is used to obtain the first server processing time based on the time when the first server receives the probe response message, the probe start time, the network transmission delay, and the network card queuing processing time.
[0121] In one optional embodiment, the probe request message includes the probe request message sending time and the probe request message receiving time. The probe request message sending time is filled in the probe request message when the first network card sends the probe request message, and the probe request message receiving time is filled in the probe request message when the second network card receives the probe request message.
[0122] In one optional embodiment, receiving the probe request message sent by the first network card includes: receiving the probe request message sent by the first network card based on the determined queue of the extended reliable datagram type and the end-to-end link; the queue of the extended reliable datagram type and the end-to-end link are determined by the first network card based on the second server information of the second server to be probed.
[0123] In one of the optional embodiments, the network information includes at least one of the network card queuing processing time of the second network card, the network transmission delay, and the first server processing time; the change of the network card queuing processing time of the second network card is used to characterize the congestion condition of the second network card; the network transmission delay is used to determine the congestion condition of the network switch between the first network card and the second network card; the first server processing time is used to determine the load condition of the first server.
[0124] Among them, the processing process of the second network card can refer to the description above and will not be elaborated here.
[0125] It should be understood that although the steps in the flowcharts involved in the above embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0126] Based on the same inventive concept, the embodiments of the present application also provide a network information detection device for implementing the network information detection method involved above. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the network information detection device provided below can refer to the limitations on the network information detection method above and will not be elaborated here.
[0127] In an exemplary embodiment, as Figure 11 shown, a network information detection device is provided, including: a detection request message generation module 1101, a first message sending module 1102, and a first message receiving module 1103, where:
[0128] The detection request message generation module 1101 is used to generate a detection request message, and the detection request message carries a detection start time;
[0129] The first message sending module 1102 is used to schedule the first network card to transmit the detection request message to the second network card, and during the transmission process of the detection request message, the detection request message sending time and the detection request message receiving time are recorded in the detection request message;
[0130] The first message receiving module 1103 is configured to receive a probe response message corresponding to a probe request message. The probe response message is transmitted by a second network card to a first server via a first network card, and during the transmission of the probe response message, network information determined based on the probe start time, the probe request message sending time, and the probe request message receiving time recorded in the probe request message is recorded in the probe response message.
[0131] In one optional embodiment, the above-mentioned first message sending module 1102 is specifically configured to schedule the first network card to send a probe request message to the second network card, and fill in the probe request message sending time in the probe request message; the probe request message is used to instruct the second network card to fill in the probe message receiving time in the probe request message after receiving the probe request message.
[0132] In one optional embodiment, the above-mentioned first message receiving module 1103 is specifically configured to receive the probe response message through the first network card. The probe response message received by the first network card carries the network card queuing processing time of the second network card obtained based on the probe request message receiving time, and the first network card obtains the network transmission delay based on the probe response message receiving time, the network card queuing processing time, and the probe request message sending time, and fills the network transmission delay into the probe response message; based on the time when the first server receives the probe response message, the probe start time, the network transmission delay, and the network card queuing processing time, the first server processing time is obtained.
[0133] In one optional embodiment, the above-mentioned probe request message generating module 1101 is specifically configured to determine the second server information based on the list of servers to be probed; generate a probe request message based on the second server information, and record the probe start time.
[0134] In one optional embodiment, the above-mentioned probe request message generating module 1101 is specifically configured to determine the queue of the extended reliable datagram type and the end-to-end link based on the second server information of the second server to be probed; the above-mentioned first message sending module 1102 is specifically configured to send doorbell information to the first network card, and the doorbell information is used to instruct the first network card to transmit the probe request message to the second network card based on the determined queue of the extended reliable datagram type and the end-to-end link.
[0135] In one alternative embodiment, the network information includes at least one of the network card queuing processing time of the second network card, the network transmission delay, and the first server processing time; the apparatus further includes: a network evaluation module configured to determine a network anomaly based on at least one of the following: the congestion condition of the second network card obtained based on the change condition of the network card queuing processing time of the second network card; the congestion condition of the network switch between the first network card and the second network card determined based on the network transmission delay; and the load condition of the first server determined based on the first server processing time.
[0136] In an exemplary embodiment, as Figure 12 shown, there is provided a network information detection apparatus, including: a second message receiving module 1201 and a second message sending module 1202, where:
[0137] The second message receiving module 1201 is configured to receive a detection request message sent by the first network card. The detection request message is generated by the first server, and the detection request message carries a detection start time. During the transmission of the detection request message, the detection request message sending time and the detection request message receiving time are recorded in the detection request message;
[0138] The second message sending module 1202 is configured to generate a detection response message corresponding to the detection request message, and transmit the detection response message to the first server via the first network card. During the transmission of the detection response message, the network information determined based on the detection start time, the detection request message sending time, and the detection request message receiving time recorded in the detection request message is recorded in the detection response message.
[0139] In one alternative embodiment, the second message sending module 1202 is specifically configured to construct an empty detection response message, copy the detection start time and the detection request message sending time in the detection request message to the empty detection response message, obtain the network card queuing processing time of the second network card based on the detection request message receiving time, and fill the network card queuing processing time into the detection response message; transmit the detection response message to the first server via the first network card. In the case where the first network card receives the detection response message, obtain the network transmission delay based on the detection response message receiving time, the network card queuing processing time, and the detection request message sending time, and fill the network transmission delay into the detection response message, and transmit the detection response message filled with the network transmission delay to the first server. The first server is configured to obtain the first server processing time based on the time when the first server receives the detection response message, the detection start time, the network transmission delay, and the network card queuing processing time.
[0140] In one optional embodiment, the probe request message includes the probe request message sending time and the probe request message receiving time. The probe request message sending time is filled in the probe request message by the first network card when sending the probe request message, and the probe request message receiving time is filled in the probe request message by the second network card when receiving the probe request message.
[0141] In one optional embodiment, the second message receiving module 1201 is specifically configured to receive a probe request message sent by the first network card based on a determined queue of the extended reliable datagram type and an end-to-end link; the queue of the extended reliable datagram type and the end-to-end link are determined by the first network card based on the second server information of the second server to be probed.
[0142] In one optional embodiment, the network information includes at least one of the network card queuing processing time of the second network card, the network transmission delay, and the first server processing time; the change situation of the network card queuing processing time of the second network card is used to characterize the congestion situation of the second network card; the network transmission delay is used to determine the congestion situation of the network switch between the first network card and the second network card; the first server processing time is used to determine the load situation of the first server.
[0143] Each module in the above network information detection device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0144] In one embodiment, a chip is provided. The chip includes a memory and a processor, and the memory stores a computer program. When the processor executes the computer program, the above access control method is implemented. Among them, the chip can be a Data Processing Unit (DPU) chip.
[0145] In one embodiment, a network interface card is provided. The network interface card includes the above chip and multiple interfaces, and the chip communicates externally through the interfaces. The interfaces include PCI / PCIE interfaces, network interfaces, etc.
[0146] In an exemplary embodiment, a computer device is provided. The electronic device includes a processor and the above network interface card. The network interface card is used to schedule messages to the processor or itself for processing, and the processor is used to process the messages scheduled by the network interface card. The computer device can be a server, and its internal structure diagram can be as Figure 13As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a network interface card. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the network interface card is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store corresponding processing data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. The computer program, when executed by the processor, implements a network information detection method.
[0147] Those skilled in the art can understand that Figure 13 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0148] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0149] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0150] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0151] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0152] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, Resistive Random Access Memory (ReRAM), Magnetoresistive Random Access Memory (MRAM), Ferroelectric Random Access Memory (FRAM), Phase Change Memory (PCM), graphene memory, etc. Volatile memory can include Random Access Memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, Artificial Intelligence (AI) processors, etc., without limitation.
[0153] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope recorded in the present application.
[0154] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A network information detection method, characterized in that: Applied to a first server, the method comprises: Generate a probe request message, wherein the probe request message carries a probe start time; Scheduling the first network card to transmit the probe request message to the second network card, and during the transmission of the probe request message, the sending time of the probe request message and the receiving time of the probe request message are recorded in the probe request message; Receive a probe response message corresponding to the probe request message, wherein the probe response message is transmitted from the second network card to the first server via the first network card, and during the transmission of the probe response message, network information determined based on the probe start time, the probe request message sending time and the probe request message receiving time recorded in the probe request message is recorded in the probe response message.
2. The method according to claim 1, characterized in that The scheduling of the first network card to transmit the probe request message to the second network card, and during the transmission of the probe request message, recording the sending time and the receiving time of the probe request message into the probe request message, includes: The first network card is scheduled to send a probe request message to the second network card, and the probe request message sending time is filled in the probe request message; the probe request message is used to instruct the second network card to fill in the probe message receiving time in the probe request message after receiving the probe request message.
3. The method according to claim 1, characterized in that The receiving a probe response message corresponding to the probe request message includes: receiving the probe response message through the first network card, wherein the probe response message received by the first network card carries the network card queue processing time of the second network card obtained based on the time when the probe request message was received, and the first network card obtains the network transmission delay based on the time when the probe response message was received, the network card queue processing time and the time when the probe request message was sent, and fills the network transmission delay into the probe response message; The first server processing time is obtained based on the time when the first server receives the detection response message, the detection start time, the network transmission delay and the network card queue processing time.
4. The method according to any one of claims 1 to 3, characterized in that: The generating of the probe request message comprises: Determine the second server information based on the list of server information to be detected; A detection request message is generated based on the second server information, and the detection start time is recorded.
5. The method according to any one of claims 1 to 3, characterized in that: Before scheduling the first network card to transmit the detection request message to the second network card, the method further includes: Determining a queue and an end-to-end link of an extended reliable datagram type based on second server information of a second server to be detected; The scheduling the first network card to transmit the detection request message to the second network card includes: Sending doorbell information to the first network card, wherein the doorbell information is used to instruct the first network card to transmit the probe request message to the second network card based on the determined queue of the extended reliable datagram type and the end-to-end link.
6. The method according to any one of claims 1 to 3, characterized in that: The network information includes at least one of the network card queue processing time of the second network card, the network transmission delay and the first server processing time; the method further includes: Determine network anomalies based on at least one of the following: Obtaining the congestion status of the second network card based on the change of the network card queue processing time of the second network card; Determine the congestion status of the network switch between the first network card and the second network card based on the network transmission delay; and A load condition of the first server is determined based on the processing time of the first server.
7. A network information detection method, characterized in that: Applied to the second network card, the method includes: Receive a probe request message sent by the first network card, where the probe request message is generated by the first server and carries a detection start time. During the transmission of the probe request message, a probe request message sending time and a probe request message receiving time are recorded in the probe request message; A probe response message corresponding to the probe request message is generated, and the probe response message is transmitted to the first server via the first network card. During the transmission of the probe response message, network information determined based on the probe start time, the probe request message sending time and the probe request message receiving time recorded in the probe request message is recorded in the probe response message.
8. The method according to claim 7, characterized in that The generating a probe response message corresponding to the probe request message, and transmitting the probe response message to the first server via the first network card, comprises: Construct an empty probe response message, copy the probe start time and the probe request message sending time in the probe request message to the empty probe response message, obtain the network card queue processing time of the second network card based on the probe request message receiving time, and fill the network card queue processing time into the probe response message; The probe response message is transmitted to the first server via the first network card. When the first network card receives the probe response message, the network transmission delay is obtained based on the reception time of the probe response message, the network card queue processing time and the probe request message sending time, and the network transmission delay is filled in the probe response message. The probe response message filled in with the network transmission delay is transmitted to the first server, and the first server is used to obtain the first server processing time based on the time when the first server receives the probe response message, the probe start time, the network transmission delay and the network card queue processing time.
9. The method according to claim 7, characterized in that: The probe request message includes a probe request message sending time and a probe request message receiving time. The probe request message sending time is filled in the probe request message when the first network card sends the probe request message, and the probe request message receiving time is filled in the probe request message when the second network card receives the probe request message.
10. The method according to any one of claims 7 to 9, characterized in that: The receiving a detection request message sent by the first network card includes: Receive a detection request message sent by the first network card based on a determined queue of an extended reliable datagram type and an end-to-end link; the queue of the extended reliable datagram type and the end-to-end link are determined by the first network card based on second server information of a second server to be detected.
11. The method according to any one of claims 7 to 9, characterized in that: The network information includes at least one of the network card queue processing time of the second network card, the network transmission delay and the first server processing time; the change of the network card queue processing time of the second network card is used to characterize the congestion of the second network card; the network transmission delay is used to determine the congestion of the network switch between the first network card and the second network card; the first server processing time is used to determine the load of the first server.
12. A network information detection device, characterized in that: The device comprises: A detection request message generation module, used to generate a detection request message, wherein the detection request message carries a detection start time; A first message sending module, used for scheduling the first network card to transmit the detection request message to the second network card, and during the transmission process of the detection request message, the detection request message sending time and the detection request message receiving time are recorded in the detection request message; The first message receiving module is used to receive a probe response message corresponding to the probe request message, wherein the probe response message is transmitted from the second network card to the first server via the first network card, and during the transmission of the probe response message, network information determined based on the probe start time, the probe request message sending time and the probe request message receiving time recorded in the probe request message is recorded in the probe response message.
13. A network information detection device, characterized in that: The device comprises: A second message receiving module is used to receive a detection request message sent by the first network card, wherein the detection request message is generated by the first server, and the detection request message carries a detection start time, and the detection request message sending time and the detection request message receiving time are recorded in the detection request message during the transmission process of the detection request message; The second message sending module is used to generate a probe response message corresponding to the probe request message, and transmit the probe response message to the first server via the first network card. During the transmission of the probe response message, network information determined based on the probe start time, the probe request message sending time and the probe request message receiving time recorded in the probe request message is recorded in the probe response message.
14. A chip comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.
15. A network interface card, characterized in that: It comprises the chip as claimed in claim 14 and a plurality of interfaces, and the chip processes data or communicates externally through the interfaces.
16. A computer device, characterized in that: The network interface card comprises the network interface card as claimed in claim 15, wherein the network interface card is used for processing data or external communication.
17. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.
18. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.