An in-band measurement method, device, electronic device and storage medium
By sending probe messages to the first satellite in the satellite network and generating the second probe messages using reinforcement learning models, the problem of insufficient flexibility of the in-band measurement scheme in the prior art is solved, and efficient and dynamic measurement of the satellite network is achieved.
Patent Information
- Application Number
- CN202510372169.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-03-27
AI Technical Summary
In the prior art, the in-band measurement scheme has poor flexibility for satellite nodes, and it is difficult to dynamically adjust the transmission of probe messages according to the network status, resulting in insufficient measurement flexibility and targeting.
By sending a first probe message to the first satellite in the satellite network, it obtains its network status information, and generates a second probe message based on the reinforcement learning model, dynamically adjusts the detection path and frequency, and sends a second probe message to the second satellite to obtain its network status information.
It improves the flexibility and pertinence of in-band measurements, can dynamically adjust the transmission of probe messages according to network status, and enhances the measurement capabilities of satellite networks.
Smart Images

Figure CN119906475B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and in particular, to an in-band measurement method, apparatus, electronic device, and storage medium. Background Art
[0002] In-band measurement is a hybrid measurement method that can measure network topology and network performance more precisely. By inserting metadata into data packets sequentially through intermediate switching nodes on the path, the acquisition of network status is realized. An in-band network telemetry system is jointly composed of a telemetry server and a switch with in-band network telemetry capabilities. The data packet processing flow of in-band network telemetry is as follows:
[0003] When a data packet arrives at a switching node, the in-band network telemetry module mirrors the packet and inserts an in-band network telemetry (INT) header according to the needs of the telemetry task; when the packet is forwarded to the last hop of the in-band network telemetry system, the switching device matches the INT header and extracts all telemetry information; the telemetry server parses the telemetry information in the telemetry packet and reports it to the upper-layer telemetry application for network monitoring, performance analysis, and fault location.
[0004] In related technologies, in-band measurement solutions generally send fixed probe packets to satellite nodes regularly to implement in-band measurement of satellite nodes, and the flexibility of in-band measurement is poor. Summary of the Invention
[0005] This application provides an in-band measurement method, apparatus, electronic device, and storage medium to solve the problem of poor flexibility of in-band measurement solutions in related technologies.
[0006] In a first aspect, this application provides an in-band measurement method, the method includes:
[0007] Sending a first probe packet to at least one first satellite to obtain first network status information of the at least one first satellite;
[0008] Generating second probe packets corresponding to at least one second satellite based on the first network status information of the at least one first satellite;
[0009] Sending the corresponding second probe packets to the at least one second satellite to obtain second network status information of the at least one second satellite.
[0010] In an optional implementation manner, the method includes:
[0011] Inputting the first network status information of the at least one first satellite into a reinforcement learning model, and generating second probe packets corresponding to at least one second satellite based on the reinforcement learning model.
[0012] In an alternative embodiment, the method further includes:
[0013] Inputting the first network state information of the at least one first satellite into a reinforcement learning model, and determining at least one detection path based on the reinforcement learning model;
[0014] The method includes:
[0015] Sending corresponding second probe messages to at least one second satellite in the at least one detection path.
[0016] In an alternative embodiment, the method includes:
[0017] Sending a first probe message to the at least one first satellite based on a first frequency;
[0018] Sending corresponding second probe messages to the at least one second satellite based on a second frequency;
[0019] Wherein, the second frequency is less than the first frequency.
[0020] In an alternative embodiment, the method further includes:
[0021] Inputting the first network state information of the at least one first satellite into a reinforcement learning model, and determining the second frequency corresponding to the at least one second satellite based on the reinforcement learning model.
[0022] In an alternative embodiment, the method includes:
[0023] Determining bitmap information corresponding to the network state information of the at least one second satellite based on the reinforcement learning model;
[0024] Sending corresponding second probe messages carrying the bitmap information to the at least one second satellite; wherein, the bitmap information is used to indicate the measured network state information.
[0025] In an alternative embodiment, the sending the first probe message to the at least one first satellite includes:
[0026] For the at least one first satellite, determining a target gateway station closest to the first satellite;
[0027] Sending the first probe message to the first satellite through the target gateway station.
[0028] In an alternative embodiment, obtaining the first network state information of the at least one first satellite includes:
[0029] For the at least one first satellite, obtain first network state information of at least one first satellite on the path from the target gateway station to the first satellite.
[0030] In an alternative embodiment, obtaining second network state information of the at least one second satellite includes:
[0031] For the at least one detection path, determine second network state information of at least one second satellite on the detection path; wherein, the second satellites on the detection path have the same second frequency.
[0032] In an alternative embodiment, determining at least one detection path based on the reinforcement learning model includes:
[0033] The reinforcement learning model is used to determine the next-hop satellite of a satellite, and the at least one detection path includes the satellite and the next-hop satellite; wherein, the reinforcement learning model is used to determine, for the satellite, the conditional probability of the satellite and at least one candidate satellite for the next-hop connection; and the candidate satellite with the maximum conditional probability is determined as the next-hop satellite of the satellite.
[0034] In an alternative embodiment, the process of determining at least one candidate satellite for the next-hop connection includes:
[0035] According to the first network state information of the satellite, determine at least one satellite adjacent to the satellite and in normal communication, and determine the at least one satellite as at least one candidate satellite.
[0036] In an alternative embodiment, the training process of the reinforcement learning model includes:
[0037] For at least one first satellite in at least one sample path, determine a first reward value corresponding to the first satellite according to the bitmap information and weight value corresponding to the network state information of the first satellite in the current iterative training process, and the sending frequency of the probe packet; determine a second reward value corresponding to the sample path according to the first reward values corresponding to at least one first satellite in the sample path; update the gradient of the reinforcement learning model, the bitmap information and weight value corresponding to the network state information, and the sending frequency of the probe packet according to the minimum second reward value corresponding to the at least one sample path; repeat the iterative training until the reinforcement learning model is trained.
[0038] In an alternative embodiment, determining the first reward value corresponding to the first satellite according to the bitmap information and weight value corresponding to the network state information of the first satellite in the current iterative process, and the sending frequency of the probe packet includes:
[0039] Determine the first sub-reward value and the resource information occupied by the probe message according to the bitmap information and weight value corresponding to the network status information of the first satellite in the current iteration process;
[0040] Determine the second sub-reward value according to the resource information occupied by the probe message and the sending frequency of the probe message;
[0041] Determine the first reward value corresponding to the first satellite according to the first sub-reward value and the second sub-reward value.
[0042] In an alternative embodiment, the method further includes:
[0043] Determine the corresponding associated network status information according to the telemetry requirements of in-band measurement;
[0044] Determine the initial bitmap information and initial weight value corresponding to the network status information of the first satellite according to the associated network status information.
[0045] In a second aspect, the present application provides an in-band measurement device, and the device includes:
[0046] A first acquisition module, configured to send a first probe message to at least one first satellite, and acquire the first network status information of the at least one first satellite;
[0047] A determination module, configured to generate a second probe message corresponding to at least one second satellite based on the first network status information of the at least one first satellite;
[0048] A second acquisition module, configured to send the corresponding second probe message to the at least one second satellite, and acquire the second network status information of the at least one second satellite.
[0049] In an alternative embodiment, the determination module is specifically configured to input the first network status information of the at least one first satellite into a reinforcement learning model, and generate a second probe message corresponding to at least one second satellite based on the reinforcement learning model.
[0050] In an alternative embodiment, the determination module is further configured to input the first network status information of the at least one first satellite into a reinforcement learning model, and determine at least one detection path based on the reinforcement learning model;
[0051] The determination module is specifically configured to send the corresponding second probe message to at least one second satellite in at least one of the at least one detection paths.
[0052] In an alternative embodiment, the first acquisition module is specifically configured to send the first probe message to the at least one first satellite based on a first frequency;
[0053] The second obtaining module is specifically configured to send corresponding second probe messages to the at least one second satellite based on a second frequency;
[0054] Wherein, the second frequency is less than the first frequency.
[0055] In an alternative embodiment, the determining module is further configured to input the first network state information of the at least one first satellite into a reinforcement learning model, and determine the second frequency corresponding to the at least one second satellite based on the reinforcement learning model.
[0056] In an alternative embodiment, the determining module is further configured to determine bitmap information corresponding to the network state information of the at least one second satellite based on the reinforcement learning model;
[0057] The second obtaining module is specifically configured to send corresponding second probe messages carrying the bitmap information to the at least one second satellite; wherein, the bitmap information is used to indicate measured network state information.
[0058] In an alternative embodiment, the first obtaining module is specifically configured to, for the at least one first satellite, determine a target gateway station closest to the first satellite; and send the first probe message to the first satellite through the target gateway station.
[0059] In an alternative embodiment, the first obtaining module is specifically configured to, for the at least one first satellite, obtain the first network state information of at least one first satellite on the path from the target gateway station to the first satellite.
[0060] In an alternative embodiment, the second obtaining module is specifically configured to, for the at least one detection path, determine the second network state information of at least one second satellite on the detection path; wherein, the second satellites on the detection path have the same second frequency.
[0061] In an alternative embodiment, the reinforcement learning model is used to determine the next-hop satellite of a satellite, and the at least one detection path includes the satellite and the next-hop satellite; wherein, the reinforcement learning model is configured to, for the satellite, determine the conditional probability of the satellite and at least one candidate satellite for the next-hop connection; and determine the candidate satellite with the maximum conditional probability as the next-hop satellite of the satellite.
[0062] In an alternative embodiment, the determining module is specifically configured to, according to the first network state information of the satellite, determine at least one satellite adjacent to the satellite and in normal communication, and determine the at least one satellite as at least one candidate satellite.
[0063] In an alternative embodiment, the device further includes:
[0064] A training module, configured to determine, for at least one first satellite in at least one sample path, a first reward value corresponding to the first satellite according to bitmap information and weight values corresponding to the network state information of the first satellite in the current iterative training process, and the sending frequency of probe messages; determine a second reward value corresponding to the sample path according to the first reward values corresponding to at least one first satellite in the sample path; update the gradient of the reinforcement learning model, the bitmap information and weight values corresponding to the network state information, and the sending frequency of probe messages according to the minimum second reward value corresponding to the at least one sample path; repeat iterative training until the reinforcement learning model is trained.
[0065] In an alternative embodiment, the training module is specifically configured to determine a first sub-reward value and resource information occupied by probe messages according to bitmap information and weight values corresponding to the network state information of the first satellite in the current iterative process; determine a second sub-reward value according to the resource information occupied by the probe messages and the sending frequency of the probe messages; determine a first reward value corresponding to the first satellite according to the first sub-reward value and the second sub-reward value.
[0066] In an alternative embodiment, the training module is further configured to determine corresponding associated network state information according to the telemetry requirements of in-band measurement; determine initial bitmap information and initial weight values corresponding to the network state information of the first satellite according to the associated network state information.
[0067] In a third aspect, the present application provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, where the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0068] The memory is used to store a computer program;
[0069] The processor is configured to implement the method when executing the program stored in the memory.
[0070] In a fourth aspect, the present application provides a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium, and the computer program implements the method when executed by a processor.
[0071] In a fifth aspect, the present application provides a computer program product, where the computer program product includes an executable program, and the executable program implements the method when executed by a processor.
[0072] In this application, first, a first probe message is sent to at least one first satellite in the satellite network, and first network status information of at least one first satellite is obtained based on the first probe message. The first network status information includes performance metric information such as satellite load, queue length, congestion status, port packet loss rate, port timestamp, port packet count, etc. Then, based on the first network status information of at least one first satellite, a second probe message corresponding to at least one second satellite is generated. Furthermore, the corresponding second probe message is sent to at least one second satellite to obtain second network status information of at least one second satellite. This application provides a solution for in-band measurement by combining two different probe messages, and the second probe message is generated based on the first network status information of at least one first satellite obtained from the first probe message. Compared with the related technology of periodically sending fixed probe messages to satellite nodes to implement in-band measurement of satellite nodes, the flexibility and pertinence of in-band measurement are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] To more clearly illustrate the technical solutions in the embodiments of this application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0074] Figure 1 Schematic diagram of the first in-band measurement process provided by this application;
[0075] Figure 2 Schematic diagram of the second in-band measurement process provided by this application;
[0076] Figure 3 Schematic diagram of the third in-band measurement process provided by this application;
[0077] Figure 4 Schematic diagram of the training process of the reinforcement learning model provided by this application;
[0078] Figure 5 Schematic diagram of the in-band measurement framework for a giant constellation network based on deep learning provided by this application;
[0079] Figure 6 Format diagram of the basic probe message provided by this application;
[0080] Figure 7 Format diagram of the dynamic probe message provided by this application;
[0081] Figure 8 Schematic diagram of the deep reinforcement learning model provided by this application;
[0082] Figure 9 Structural schematic diagram of the in-band measurement device provided by this application;
[0083] Figure 10 Structural schematic diagram of the electronic device provided by this application. Detailed implementation manners
[0084] To make the purpose and implementation manners of this application clearer, the following will clearly and completely describe the exemplary implementation manners of this application in combination with the drawings in the exemplary embodiments of this application. Obviously, the described exemplary embodiments are only a part of the embodiments of this application, rather than all the embodiments.
[0085] It should be noted that the brief description of the terms in this application is only for the convenience of understanding the subsequent described implementation manners, rather than intending to limit the implementation manners of this application. Unless otherwise specified, these terms should be understood in their ordinary and common meanings.
[0086] The terms "first", "second", "third", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar or homogeneous objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms can be interchanged under appropriate circumstances.
[0087] The terms "comprising" and "having" and any variations thereof are intended to cover but not exclude inclusion. For example, a product or device comprising a series of components does not necessarily have to be limited to all the components clearly listed, but may include other components not clearly listed or inherent to these products or devices.
[0088] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or a combination of hardware or / and software code that can perform functions related to this element.
[0089] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of this application.
[0090] For the sake of convenience in explanation, the above description has been made in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. According to the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are for better explaining the principles and practical applications, so that those skilled in the art can better use the embodiments and various different modified embodiments suitable for specific use considerations.
[0091] In the related art, the data packet processing flow of in-band network telemetry is as follows:
[0092] 1. When an ordinary data packet arrives at the first switching node of the in-band network telemetry system, the in-band network telemetry module matches and mirrors the packet through the sampling method set on the switch, inserts the in-band network telemetry INT (In-band Network Telemetry) header after the four-layer header according to the needs of the telemetry task, and encapsulates the telemetry information specified by the INT header into metadata (MetaData, MD) and inserts it after the INT header;
[0093] 2. When the packet is forwarded to an intermediate node, the device matches the INT header and inserts MD;
[0094] 3. When the packet is forwarded to the last hop of the in-band network telemetry system, the switching device matches the INT header, inserts the last MD, extracts all the telemetry information, and forwards it to the telemetry server through methods such as g Remote Procedure Call (gRPC);
[0095] 4. The telemetry server parses the telemetry information in the telemetry packet, reports it to the upper-layer telemetry application program for network monitoring, performance analysis, and fault location.
[0096] In the related technical solutions, network telemetry is mainly divided into two categories: passive network telemetry (PNT) and active network telemetry (ANT). Passive network telemetry technology relies on network devices to automatically add internal information, such as link utilization and queuing delay, during the data packet transmission process for centralized controller analysis. Active network telemetry technology, on the other hand, actively generates and sends telemetry probes through source routing technology to detect user-specified network paths. The ANT system in the related art can optimize and flexibly plan the probe path based on a fixed network environment and performance metrics to meet different telemetry requirements.
[0097] In the related art, the active network telemetry method lacks the ability to quickly adapt to new telemetry requirements and self-adjust in a dynamic network environment; although the passive network telemetry method provides fine-grained network status data, it is difficult to achieve full network coverage and may increase network overhead due to the uncontrolled probe path. Its measurement means mainly include two types: periodically sending network-wide probes and manually setting probe parameters, with weak adaptability. In the method of periodically sending network-wide probes, if the probe sending period is too large, the network measurement efficiency will be reduced, and network problems cannot be detected in time; if the period is too small, the network overhead will increase sharply, and it cannot send probes locally according to the actual state of the constellation network, with poor flexibility. Compared with the method of periodically sending network-wide probes, manually setting probe parameters provides operational flexibility, but its effectiveness highly depends on the experience accumulation and subjective judgment of measurement personnel. It often requires multiple trial-and-error and adjustment processes. When the application scenario expands to a large-scale constellation network environment, the complexity and dynamicity of the system increase sharply, making the pre-set parameters and rules difficult to fully adapt to various changing situations; in practice, if the threshold is set too high, key features in the signal may be missed, and if set too low, the result may be affected by too much noise and interference, reducing the measurement accuracy.
[0098] Traditional in-band measurement methods are mainly proposed for terrestrial mobile networks, with relatively simple scenarios, and the performance indicators that can be measured cannot meet the indicator requirements of satellite communication networks such as bit error rate. At the same time, it cannot adaptively adjust measurement indicators according to measurement requirements.
[0099] In addition, in the related art, network fault location based on in-band measurement mainly relies on periodically sending network-wide probes, with high overhead and poor flexibility. At the same time, it also lacks an autonomous learning mechanism and cannot fully explore the deep information of historical measurement data to dynamically guide network fault discovery and location.
[0100] To solve the above technical problems, this application provides an in-band measurement scheme that combines static probe messages and dynamic probe messages. Figure 1 The following is the first schematic diagram of the in-band measurement process provided by this application, including the following steps:
[0101] S101: Send a first probe message to at least one first satellite to obtain the first network status information of the at least one first satellite;
[0102] S102: Generate a second probe message corresponding to at least one second satellite based on the first network status information of the at least one first satellite;
[0103] S103: Send the corresponding second probe message to the at least one second satellite to obtain the second network status information of the at least one second satellite.
[0104] The in-band measurement method provided by this application is applied to an electronic device, which can be a network-side device such as an eNB (Evolved NodeB, radio base station) on the network side; it can also be a terminal device such as a mobile phone or a computer; it can also be a server.
[0105] For multiple satellites in a satellite network, the electronic device sends a first probe message to at least one first satellite in the satellite network. Optionally, the electronic device can send the first probe message to all first satellites in the satellite network; or according to the satellite domain specified in the telemetry requirement, send the first probe message to all first satellites in the specified satellite domain of the satellite network; or according to at least one satellite specified in the telemetry requirement, send the first probe message to the at least one satellite. The first probe message is also called a basic probe message. The electronic device sends the first probe message to at least one first satellite to obtain the first network status information of the at least one first satellite. Optionally, the first network status information includes performance index information such as satellite load, queue length, congestion status, port packet loss rate, port timestamp, and port packet count. In addition, the first network status information can also include the topology information of the satellite network. According to the topology information of the satellite network, for each first satellite, the adjacent satellites that communicate normally with the first satellite can be determined.
[0106] After sending the first probe message to at least one first satellite and obtaining the first network status information of the at least one first satellite, then based on the first network status information of the at least one first satellite, generate second probe messages corresponding to at least one second satellite. The second probe message can be understood as a dynamic probe message. Since the first network status information of the at least one first satellite is different, the second probe messages generated corresponding to the at least one second satellite may also be different. Then send the corresponding second probe messages to at least one second satellite, and for the at least one second satellite, based on the second probe message corresponding to the second satellite, obtain the second network status information of the second satellite.
[0107] In this application, first, a first probe message is sent to at least one first satellite in the satellite network, and based on the first probe message, the first network status information of at least one first satellite is obtained. The first network status information includes performance metric information such as satellite load, queue length, congestion status, port packet loss rate, port timestamp, port packet count, etc. Then, based on the first network status information of at least one first satellite, a second probe message corresponding to at least one second satellite is generated. Furthermore, the corresponding second probe message is sent to at least one second satellite to obtain the second network status information of at least one second satellite. This application provides a solution for in-band measurement by combining two different probe messages, and the second probe message is generated based on the first network status information of at least one first satellite obtained from the first probe message. Compared with the related technology of periodically sending fixed probe messages to satellite nodes to implement in-band measurement of satellite nodes, the flexibility and pertinence of in-band measurement are improved.
[0108] In an alternative embodiment, the method includes:
[0109] Input the first network status information of the at least one first satellite into a reinforcement learning model, and based on the reinforcement learning model, generate a second probe message corresponding to at least one second satellite.
[0110] A pre-trained reinforcement learning model is deployed in the electronic device. It should be noted that the reinforcement learning model can be trained in the main electronic device that executes the in-band measurement method; or the reinforcement learning model is trained in other electronic devices and then the trained reinforcement learning model is deployed to the main electronic device that executes the in-band measurement method. The reinforcement learning model is used to generate a second probe message corresponding to at least one second satellite. Optionally, after determining the first network status information of at least one first satellite, input the first network status information of at least one first satellite into the trained reinforcement learning model, and based on the reinforcement learning model, generate a second probe message corresponding to at least one second satellite.
[0111] In an alternative embodiment, the method further includes:
[0112] Input the first network status information of the at least one first satellite into a reinforcement learning model, and determine at least one detection path based on the reinforcement learning model;
[0113] The method includes:
[0114] Send the corresponding second probe message to at least one second satellite in the at least one detection path.
[0115] The reinforcement learning model is also used to predict the probing path of in-band measurement. Optionally, in each first satellite, a second satellite is randomly initialized, and then the first network state information of the second satellite and the first network state information of each satellite that is one-hop connected to the second satellite are input into the reinforcement learning model. Based on the reinforcement learning model, the next-hop second satellite of the initialized second satellite is selected from each satellite that is one-hop connected to the second satellite. Then, the first network state information of the next-hop second satellite and the first network state information of each satellite that is one-hop connected to the next-hop second satellite are input into the reinforcement learning model. Based on the reinforcement learning model, the next-next-hop second satellite of the next-hop second satellite is selected from each satellite that is one-hop connected to the next-hop second satellite. And so on, until the number of second satellites determined in the probing path reaches a preset number threshold, and then this probing path is obtained.
[0116] Then, a second satellite is randomly initialized from the first satellites other than the second satellites included in the determined probing path, and the above steps are continued to repeat, and another probing path is obtained. In this way, at least one probing path can be determined based on the reinforcement learning model. The number of satellites included in the probing path does not exceed the preset number threshold, and the probing direction of the probing path is the next-hop connection direction between the second satellites.
[0117] For example, a certain probing path includes 5 second satellites, namely second satellite A, second satellite B, second satellite C, second satellite D, and second satellite E. Second satellite B is the next-hop satellite of second satellite A, second satellite C is the next-hop satellite of second satellite B, second satellite D is the next-hop satellite of second satellite C, and second satellite E is the next-hop satellite of second satellite D. Then, the probing direction of this probing path is from second satellite A to second satellite B to second satellite C to second satellite D to second satellite E.
[0118] Send corresponding second probe messages to at least one second satellite in at least one detection path, and obtain second network status information of at least one second satellite. Optionally, determine the first second satellite in the detection path according to the detection direction of the detection path, and then send a second probe message to the first second satellite, and obtain the second network status information of the first second satellite based on the second probe message. Optionally, when it is necessary to measure the network status information of a certain second satellite in the middle of the detection path, a second probe message can also be sent to the first second satellite first, and then the second probe message can be forwarded in sequence according to the detection direction of the detection path to the network status information of a certain second satellite to be measured, and then the second network status information of the certain second satellite can be obtained based on the second probe message. Optionally, a second probe message can also be sent to the first second satellite, and then the second probe message can be forwarded in sequence according to the detection direction of the detection path to all the second satellites included in the detection path, and then the second network status information of all the second satellites included in the detection path can be obtained. The second probe message can be understood as a dynamic probe message.
[0119] In this application, first send a first probe message to at least one first satellite in the satellite network, and obtain the first network status information of at least one first satellite based on the first probe message. The first network status information includes performance index information such as satellite load, queue length, congestion status, port packet loss rate, port timestamp, and port packet count. Then input the first network status information of at least one first satellite into a reinforcement learning model, determine at least one detection path and generate a second probe message corresponding to at least one second satellite based on the reinforcement learning model. Further, along at least one detection path, send corresponding second probe messages to at least one second satellite in at least one detection path, and obtain the second network status information of at least one second satellite. This application provides an in-band measurement scheme based on deep learning, that is, after obtaining the first network status information of at least one first satellite, based on the reinforcement learning model obtained by deep learning, determine at least one detection path, and then obtain the second network status information of at least one second satellite in at least one detection path. Compared with the related technology of periodically sending fixed probe messages to satellite nodes to implement in-band measurement of satellite nodes, the flexibility and pertinence of in-band measurement are further improved.
[0120] In an alternative embodiment, the method includes:
[0121] Send a first probe message to the at least one first satellite based on a first frequency;
[0122] Send corresponding second probe messages to the at least one second satellite based on a second frequency;
[0123] Wherein, the second frequency is less than the first frequency.
[0124] The method further includes:
[0125] Input the first network status information of the at least one first satellite into a reinforcement learning model, and determine the second frequency corresponding to the at least one second satellite based on the reinforcement learning model.
[0126] The first frequency can be a preset frequency. The second frequency can be the frequency for the second satellite corresponding to the determined second satellite to send a second probe message based on the reinforcement learning model.
[0127] Optionally, if the second network status information of each second satellite in the detection path is separately obtained, the second frequencies corresponding to each second satellite are the same or different.
[0128] If a second probe message is sent to the first second satellite, and then the second probe message is sequentially forwarded along the detection direction of the detection path to all the second satellites included in the detection path, and then the second network status information of all the second satellites included in the detection path is obtained. At this time, the second frequencies corresponding to each second satellite included in the detection path are the same.
[0129] Figure 2 The second in-band measurement process schematic diagram provided by this application includes the following steps:
[0130] S201: Send a first probe message to at least one first satellite based on the first frequency, and obtain the first network status information of the at least one first satellite;
[0131] S202: Input the first network status information of the at least one first satellite into a reinforcement learning model, determine at least one detection path based on the reinforcement learning model, and generate a second probe message corresponding to at least one second satellite;
[0132] S203: Send a second probe message to at least one second satellite in the at least one detection path based on the second frequency, and obtain the second network status information of the at least one second satellite; wherein, the second frequency corresponding to the at least one second satellite is determined based on the reinforcement learning model.
[0133] In an optional implementation manner, the method includes:
[0134] Determine bitmap information corresponding to the network status information of the at least one second satellite based on the reinforcement learning model;
[0135] Send a corresponding second probe message carrying the bitmap information to the at least one second satellite; wherein, the bitmap information is used to indicate the measured network status information.
[0136] The bitmap information is used to indicate which network status information needs to be measured. For example, the network status information includes six types of performance metric information such as satellite load, queue length, congestion status, port packet loss rate, port timestamp, and port packet count. The first to sixth bits in the bitmap information respectively represent the above six types of performance metric information. For each bit, if the value in the bitmap information is 1, it means to measure the performance metric parameter of this bit; if the value in the bitmap information is 0, it means not to measure the performance metric parameter of this bit. For example, for a certain second satellite, the bitmap information corresponding to the network status information to be measured of this second satellite is 1, 1, 1, 0, 0, 0, which means to measure the satellite load, queue length, and congestion status of this second satellite, while not measuring the port packet loss rate, port timestamp, and port packet count.
[0137] The second probe message sent by the electronic device to at least one second satellite carries the corresponding bitmap information, so as to obtain the network status information indicated by the corresponding bitmap information of this second satellite.
[0138] Figure 3 This is the schematic diagram of the third in-band measurement process provided by this application, including the following steps:
[0139] S301: Send a first probe message to at least one first satellite based on a first frequency, and obtain the first network status information of the at least one first satellite;
[0140] S302: Input the first network status information of the at least one first satellite into a reinforcement learning model, determine at least one detection path based on the reinforcement learning model, generate second probe messages corresponding to at least one second satellite, and determine the bitmap information corresponding to the network status information of at least one second satellite;
[0141] S303: Send the corresponding second probe message carrying the bitmap information to at least one second satellite in the at least one detection path based on a second frequency; obtain the second network status information of the at least one second satellite; wherein, the second frequency corresponding to the at least one second satellite is determined based on the reinforcement learning model.
[0142] In an optional implementation manner, the sending the first probe message to at least one first satellite includes:
[0143] For the at least one first satellite, determine the target gateway station closest to the first satellite; send the first probe message to the first satellite through the target gateway station.
[0144] To improve the efficiency of in-band measurement, when an electronic device sends a first probe message to at least one first satellite, for at least one first satellite, it first determines a target gateway station closest to the first satellite. The electronic device first sends the first probe message to the target gateway station, and then the target gateway station sends the first probe message to the first satellite.
[0145] In one case, for at least one first satellite, the target gateway station closest to the first satellite communicates directly with the first satellite. At this time, the electronic device first sends the first probe message to the target gateway station, and then the target gateway station sends the first probe message to the first satellite, thereby obtaining the first network status information of the first satellite. In another case, for at least one first satellite, the target gateway station closest to the first satellite communicates with the first satellite through multiple hops. At this time, the electronic device first sends the first probe message to the target gateway station, and then the target gateway station sends the first probe message to at least one intermediate routing device. Then, after being forwarded by at least one routing device, the first probe message is sent to the first satellite, thereby obtaining the first network status information of the first satellite. The at least one routing device may be another gateway station or another satellite.
[0146] In this application, for each first satellite, a target gateway station closest to the first satellite can be determined; the first probe message is sent to the first satellite through the target gateway station to obtain the first network status information of the first satellite.
[0147] Preferably, obtaining the first network status information of the at least one first satellite includes:
[0148] For the at least one first satellite, obtaining the first network status information of at least one first satellite on the path from the target gateway station to the first satellite.
[0149] In this application, for at least one first satellite, the target gateway station closest to the first satellite communicates with the first satellite through multiple hops. At this time, the electronic device first sends the first probe message to the target gateway station, and then the target gateway station sends the first probe message to at least one intermediate routing device. Then, after being forwarded by at least one routing device, the first probe message is sent to the first satellite. At this time, the path from the target gateway station to the first satellite includes at least one other first satellite. In this application, the first network status information of at least one first satellite on the path from the target gateway station to the first satellite can be obtained. Preferably, the first network status information of all first satellites on the path from the target gateway station to the first satellite can be obtained. In this way, for the intermediate first satellites whose first network status information has been determined, there is no need to determine their corresponding target gateway stations, thereby saving resources for determining the first network status information of each first satellite.
[0150] In an alternative embodiment, obtaining the second network state information of the at least one second satellite includes:
[0151] For the at least one detection path, determining the second network state information of at least one second satellite on the detection path; wherein, the second satellites on the detection path have the same second frequency.
[0152] In this application, for at least one detection path, the electronic device sends a second probe message to the first second satellite of the detection path, then forwards the second probe message to all the second satellites included in the detection path in the detection direction of the detection path, and then obtains the second network state information of at least one second satellite included in the detection path. Optionally, the second network state information of all the second satellites included in the detection path can be obtained. At this time, since the second network state information of all the second satellites included in the detection path is obtained simultaneously, the second frequencies corresponding to the respective second satellites included in the detection path are the same.
[0153] In an alternative embodiment, determining at least one detection path based on the reinforcement learning model includes:
[0154] The reinforcement learning model is used to determine the next-hop satellite of the satellite, and the at least one detection path includes the satellite and the next-hop satellite; wherein, the reinforcement learning model is used to determine the conditional probability of the satellite and at least one candidate satellite for the next-hop connection for the satellite; and the candidate satellite with the maximum conditional probability is determined as the next-hop satellite of the satellite.
[0155] Randomly initialize a satellite as the satellite to be processed, and then determine the satellite to be processed and at least one candidate satellite for the next-hop connection. Then, input the first network state information of the satellite to be processed and the first network state information of at least one candidate satellite for the next-hop connection into the reinforcement learning model, and determine the conditional probability of at least one candidate satellite for the next-hop connection based on the reinforcement learning model.
[0156] Wherein, the weight values corresponding to each performance index included in the network state information are obtained through training based on the reinforcement learning model, and then for at least one candidate satellite for the next-hop connection of the satellite to be processed, according to the formula determine the conditional probability of the candidate satellite, and then determine the candidate satellite with the maximum conditional probability as the next-hop satellite of the satellite to be processed. Wherein, the of the satellite represents n performance indexes. represents the weight values corresponding to the respective n performance indexes, obtained through training.
[0157] In this application, the process of determining candidate satellites for the at least one next-hop connection includes:
[0158] Based on the first network status information of the satellite, determine at least one satellite adjacent to and in normal communication with the satellite, and determine the at least one satellite as at least one candidate satellite.
[0159] Based on the first network status information of each first satellite, for each first satellite, the satellites in normal communication with the first satellite can be determined. Normal communication means that the communication function is normal. Adjacent means that the physical positions of the satellites are adjacent, including adjacent left and right and adjacent up and down; or it can also include diagonal adjacency.
[0160] Figure 4 The following is a schematic diagram of the training process of the reinforcement learning model provided in this application, including the following steps:
[0161] S401: For at least one first satellite in at least one sample path, based on the bitmap information and weight value corresponding to the network status information of the first satellite in the current iterative training process, and the sending frequency of the probe message, determine the first reward value corresponding to the first satellite;
[0162] S402: Based on the first reward values corresponding to at least one first satellite in the sample path, determine the second reward value corresponding to the sample path;
[0163] S403: Update the gradient of the reinforcement learning model, the bitmap information and weight value corresponding to the network status information, and the sending frequency of the probe message according to the minimum second reward value corresponding to the at least one sample path; repeat the iterative training until the reinforcement learning model is trained.
[0164] In this application, after obtaining the first network status information of each first satellite, the reinforcement learning model is trained based on the first network status information of each first satellite. When the first network status information of each first satellite is determined again next time, the reinforcement learning model is trained again based on the first network status information of each first satellite next time, and so on.
[0165] The following describes the training process of the reinforcement learning model. First, it should be noted that each time after obtaining the first network status information of each first satellite, the training of the reinforcement learning model needs to be completed through multiple iterative parameter updates.
[0166] First, according to the first network state information of each first satellite and the weight value in the current iterative training process, at least one sample path is obtained through the conditional probability calculation formula. For at least one first satellite among the at least one sample path, according to the bitmap information and weight value corresponding to the network state information of this first satellite in the current iterative training process, and the sending frequency of the dynamic probe message of this first satellite in the current iterative training process, the first reward value corresponding to this first satellite is determined.
[0167] The determining of the first reward value corresponding to this first satellite according to the bitmap information and weight value corresponding to the network state information of this first satellite in the current iterative process, and the sending frequency of the probe message includes:
[0168] According to the bitmap information and weight value corresponding to the network state information of this first satellite in the current iterative process, determine the first sub - reward value and the resource information occupied by the probe message;
[0169] According to the resource information occupied by the probe message and the sending frequency of the probe message, determine the second sub - reward value;
[0170] According to the first sub - reward value and the second sub - reward value, determine the first reward value corresponding to this first satellite.
[0171] The bitmap information and weight value corresponding to the network state information of the first satellite are different, and the resource information occupied by the probe message corresponding to the first satellite is also different. In the electronic device, the corresponding relationship between the bitmap information and weight value corresponding to the network state information and the resource information occupied by the probe message can be pre - saved. In this way, according to the bitmap information and weight value corresponding to the network state information of this first satellite in the current iterative process, the resource information occupied by the probe message is determined. In addition, according to the bitmap information and weight value corresponding to the network state information of this first satellite in the current iterative process, the first sub - reward value is determined. Optionally, the first sub - reward value can be determined according to the formula Determine the first sub - reward value.
[0172] According to the resource information occupied by the probe message and the sending frequency of the probe message, determine the second sub - reward value. Optionally, the ratio of the resource information occupied by the probe message and the sending frequency of the probe message is determined as the second sub - reward value.
[0173] According to the first sub - reward value and the second sub - reward value, determine the first reward value corresponding to this first satellite. Optionally, the sum value of the first sub - reward value and the second sub - reward value is determined as the first reward value corresponding to this first satellite.
[0174] Determine the second reward value corresponding to the sample path according to the first reward value corresponding to at least one first satellite in the sample path. Optionally, determine the sum value of the first reward values corresponding to at least one first satellite in the sample path as the second reward value corresponding to the sample path. In this way, each sample path corresponds to a second reward value, and the smallest second reward value is selected from them to update the gradient of the reinforcement learning model, the bitmap information corresponding to the network state information and the weight value, and the sending frequency of the probe message; repeat the iterative training until the reinforcement learning model is trained. Optionally, when the number of iterative training times meets the requirements, it is determined that the reinforcement learning model is trained; or, when the difference between the smallest second reward values obtained from two adjacent trainings is less than the set threshold, it is determined that the reinforcement learning model is trained.
[0175] In this application, the method further includes:
[0176] Determine the corresponding associated network state information according to the telemetry requirements of in-band measurement;
[0177] According to the associated network state information, determine the initial bitmap information and the initial weight value corresponding to the network state information of the first satellite.
[0178] The telemetry requirements of in-band measurement carry the targets of in-band measurement. For example, the targets of in-band measurement are to measure the congestion state, measure the satellite load, etc. The association relationship between different telemetry requirements and network state information is pre-stored in the electronic device. Association can be understood as the network state information that has an impact on the target of the telemetry requirement is the associated network state information. The network state information that has no impact on the target of the telemetry requirement is the non-associated network state information. The initial bitmap value corresponding to the associated network state information is 1, and the initial bitmap value corresponding to the non-associated network state information is 0. In addition, different network state information has different degrees of influence on the target of the telemetry requirement. The greater the degree of influence, the greater the initial weight value corresponding to the network state information. Based on the above strategy, according to the associated network state information, determine the initial bitmap information and the initial weight value corresponding to the network state information of the first satellite. Subsequently, update the gradient of the reinforcement learning model, the bitmap information and the weight value corresponding to the network state information, and the sending frequency of the probe message according to the determined smallest second reward value; repeat the iterative training until the reinforcement learning model is trained.
[0179] The following will describe in detail the in-band measurement method provided by this application based on deep learning with reference to the accompanying drawings.
[0180] Figure 5 This is the framework diagram of in-band measurement implemented by the giant constellation network based on deep learning provided by this application, as Figure 5As shown in the figure, this application proposes an in-band measurement method for a giant constellation network based on deep learning. This method consists of three main parts: a basic in-band measurement module, a dynamic in-band measurement module, and a telemetry analyzer module. Users can specify telemetry tasks through a network application, provide telemetry requirements to the telemetry analyzer, and distribute them to the basic in-band measurement module. The basic in-band measurement module sends basic measurement probes (the first probe message) to the satellite nodes of the constellation network to obtain basic network status information (the first network status information). This information is then used as the initial input for the dynamic in-band measurement module, which drives the dynamic in-band measurement module to generate dynamic measurement probes (the second probe message). The results of the dynamic measurement probes can also be used as input for the dynamic in-band measurement module for reinforcement learning to select satellite nodes or paths with risks for further dynamic telemetry. The results of the basic measurement probes and the dynamic measurement probes are simultaneously input into the telemetry analyzer module for network fault analysis, location, and adaptive adjustment of alarm thresholds / rules.
[0181] The performance metrics collected for the results of the dynamic probes can be divided into three categories: node level, port level, and bitmap level. At the node level, signal strength, link quality, congestion, etc. can be measured. At the port level, data transmission rate, bandwidth utilization, bit error rate, packet loss rate, latency, and jitter can be measured. The bitmap level focuses on signal strength, signal-to-noise ratio, link quality, and queue length.
[0182] The basic in-band measurement module collects basic network parameters through basic probes. This information can be used to understand the current state of the network and provides basic data for the path planning of subsequent dynamic probes. Through the basic probe path selection method, the path is reasonably planned, and the basic probes can traverse all links in the network, thus achieving full network coverage. The basic probes can also be used for the location of network faults. When a certain link fails, the fault location can be quickly located by checking which basic probes cannot reach the controller.
[0183] Figure 6 This is the format diagram of the basic probe message provided by this application, as Figure 6As shown, the head of the basic probe is similar to that of an ordinary data packet. It consists of an Ethernet header ETH, an Internet Protocol IP header, and a TCP / UDP header. TCP (Transmission Control Protocol) and UDP (User Datagram Protocol) are two main transport layer protocols in the Internet protocol suite. The source port is Source Port, and the destination port Destination Port is set to "Base_probe" to ensure that the programmable satellite routing node can resolve the basic probe. In addition, the destination address of the basic probe is set to the IP address of the controller to forward the basic probe to the controller. Additionally, the basic probe also includes a source routing stack (SR stack, i.e., the INT header) and an information stack (INFO stack, i.e., Metadata). The SR stack is a fixed-length stack composed of a series of port labels (Port). Each port label records a hop on the path of the basic probe, and the satellite routing switch node forwards the basic probe based on these port labels. The INFO stack is used to store network information, and each INFO label contains the switch ID (Swtich ID) and the value (Value) of each network metadata.
[0184] The main task of the basic probe is to collect the basic state information of the network, which is crucial for dynamic probe path planning. The basic probe is based on the shortest geodesic distance as the matching rule, and the fixed-length SR stack can ensure that it will not be interrupted due to stack capacity issues during the algorithm execution, thus guaranteeing the effectiveness and reliability of path planning. Since the collection frequency of the basic probe is relatively low and the amount of data carried is small, the fixed-length SR stack will not cause excessive consumption of network resources. If a variable-length SR stack is adopted, it may lead to unnecessary resource waste due to excessive stack capacity or be unable to meet the requirements of path planning due to insufficient stack capacity.
[0185] The basic probe path selection is used for the routing paths through which the probe is sent and received. To ensure the minimum overhead and real-time guarantee, this application uses the nearest gateway station as the path selection criterion. Set the sampling interval, traverse the entire measurement time period, calculate the geodesic distances from the target satellite node to different gateway stations at different times, and use the shortest geodesic distance as the matching rule for satellite node-gateway station matching. The probe message is sent by the network controller and directly sent to the target satellite node via the gateway station. To ensure that they can traverse the entire network topology and collect all network state information, the propagation path may involve multiple network nodes (including satellites and possible gateway stations). For the preferred gateway station, if a gateway station cannot reach a certain node, it will be forwarded through other satellites.
[0186] The matching calculation process between satellite nodes and gateway stations is as follows:
[0187]
[0188] Among them, satId is the selected satellite node, trajectory is the available path of the selected node, step is the sampling interval, with a default value of 60s, samplePoints is the sampling result of the trajectory line, gwsList is the list of gateway stations, gwId is the ID of the gateway station closest to the satellite node at the current moment, distance is the corresponding shortest geodesic distance, and satGwMap is the matching result between satellite nodes and gateway stations.
[0189] The dynamic in-band measurement module collects specific network information required by network applications through dynamic probes based on basic probes. This information includes but is not limited to node load, queue length, congestion status, port packet loss rate, port timestamp, port packet count, etc. By adjusting the path of the dynamic probe and the type of information collected, diverse telemetry requirements can be met. Compared with basic probes, dynamic probes are injected into the network at a higher frequency, so they can collect network status information more frequently. This high-frequency data collection helps to achieve more fine-grained network monitoring. The path of the dynamic probe is dynamically planned according to the real-time network status. The dynamic probe path discovery method based on deep reinforcement learning can adjust the probe path in real time according to factors such as network load and topology changes to optimize telemetry performance. By adjusting the reward function, multi-objective optimization can be supported. When the network topology or load changes, the dynamic probe can quickly adjust the path and locate faults to ensure the accuracy and timeliness of telemetry information.
[0190] Figure 7 This is the format diagram of the dynamic probe message provided by this application. As Figure 7 shown, the dynamic probe is responsible for dynamically collecting specific network information required by the application. The dynamic probe format is similar to the basic probe format. Since the length change range of the dynamic probe path is larger, using a fixed-length SR stack may cause waste. In addition, considering the diverse types of network metadata collected by the dynamic probe, two new fields are introduced in the dynamic probe format: Length and Bitmap. As Figure 7As shown, in the dynamic probe format, the SR stack is a stack with variable length, and the "length" field determines the length of the SR stack. The "bitmap" field determines which types of metadata should be collected by the INFO tag, with each bit corresponding to a type of metadata. If a bit is set to 1, it means that the DP needs to collect the corresponding type of metadata when passing through the corresponding node; if it is set to 0, it means that no collection is required. The destination port in the dynamic probe is set to "Dynamic_probe" so that the satellite routing node can parse the dynamic probe. Through the Bitmap field in the INFO stack, the system can selectively collect specific types of metadata at different ports, further improving the pertinence and efficiency of network telemetry.
[0191] Figure 8 This is a schematic diagram of the deep reinforcement learning model provided by this application. This application uses the basic network state information of the entire network collected by the basic probe for path planning to ensure that the DP can effectively cover and meet specific telemetry requirements.
[0192] Network topology: Represented as an undirected graph G=(V, E), where V is the set of nodes and E is the set of edges. Each node represents a device (such as a switch) in the network, and each edge represents the connection between devices.
[0193] Node information: Includes the connection status of each node (i.e., the connection relationship with other nodes) and port delay information, etc. This information is used to evaluate the performance of different paths.
[0194] The output of this application covers dynamic node selection, prediction of dynamic node parameter masks (bitmap Bitmap), and prediction of transmission cycles and intervals. The dynamic probe path consists of a series of node sequences, clearly defining the transmission route of the probe in the network and collecting necessary network state information when passing through each node. The Bitmap prediction generates a binary bitmap that matches the number of types of metadata that can be collected in the network. The bit values of 0 and 1 determine whether the dynamic probe needs to collect specific metadata at each node on the transmission path, achieving refined control of data collection. Finally, the transmission cycle prediction outputs a time interval value (representing the frequency), which is dynamically adjusted according to network load fluctuations, changes in telemetry requirements, and the network environment. It represents the time interval between two consecutive dynamic probe injections into the network, thus effectively managing network overhead while ensuring data timeliness.
[0195] In the dynamic path discovery method, the model adaptively selects nodes and ports (i.e., dynamic node - port selection), and determines the information metrics to be collected (i.e., metric - info bitmap selection). The model can measure network performance parameters at the node, port, and bitmap levels. The model dynamically adjusts the sending frequency of dynamic probes and the path planning strategy according to real - time network information and telemetry requirements. This adaptive ability enables the model to quickly respond to changes in the network state, optimize the sending timing and path of probes, thereby improving the accuracy and efficiency of telemetry. The model calculates an appropriate sending interval and number of sends based on the network topology, link status, and requirements of the telemetry task, and adjusts the sending interval and number of probes through bandwidth utilization. When a fault occurs in the network, the model immediately stops sending dynamic probes.
[0196] In the network telemetry system, the j - th dynamic probing path is represented as , where p j,i represents the i - th node passed by path p j , and N p is the total number of nodes in this path. The set of links covered by path N p is denoted as l p . In the network telemetry system, there are a total of P paths. For the j - th path, the nodes passed through are represented as pj. The set of links connecting these nodes is called lq. The concept of lq is only for a single dynamic probing path and serves as a symbolic identifier. When planning this path, various performance metrics need to be considered comprehensively, including signal strength, link quality, and congestion status at the node level, data transfer rate, bandwidth utilization, bit error rate, packet loss rate, latency, and jitter at the port level, and signal strength, signal - to - noise ratio, link quality, and queue length at the Bitmap level. All telemetry task performance metrics are aggregated into a set , where m i is the value of performance metric m i . Given that the importance of each performance metric varies, a weight w i is assigned to each metric m i , and the set of weights is . To ensure comprehensive coverage of the service network, the dynamic probing path planning problem can be formulated as minimizing the sum C of the product of each performance metric and its weight, which can be expressed as: .
[0197] The input set contains the connection status of node i with other nodes and the delay information of each port. X0 contains all the initial information required to generate the sequence. At each step t (from 0 to T), it is necessary to determine the next element y based on the current state (i.e., the part of the sequence that has been generated t ) and the current input set X t+1。The calculation formula of its probability sequence can be expressed as:
[0198] ;
[0199] is the conditional probability of selecting the next element y t under the condition of the known sequence Y t generated in the previous t steps and the current input set X t+1 . The calculation steps are as follows.
[0200] Initialization: Set t = 0, Y0 = 0 (empty sequence), and the initial input set X0.
[0201] X i can be understood as the network state information of the i-th node, which nodes this node is connected to; Y can be understood as the generated path, T is the length of the path, and t is each step passed. One node is determined at each step.
[0202] Recursive calculation:
[0203] For each step t (from 0 to T - 1): According to the current sequence Y t and the input set X t , calculate the conditional probability . Select a y that maximizes t+1 , and add it to the sequence, that is , to obtain the optimal choice. Update the input set X t+1 to reflect the latest sequence state.
[0204] Termination condition: When t = T, terminate the calculation. T is the number of satellites included in the preset path.
[0205] Provide the ability in two directions when performing path planning: Select nodes through the probability sequence ; Select performance by minimizing .
[0206] State: At the end of any time point, the current decision-making environment state can be expressed as s t . The state s t includes two parts of information: the current node index and the set of links for the remaining telemetry requirements. The node index has two states: state_index indexes a node in the network, and state_create or the special value 0 indicates that a new path is being considered.
[0207] Action: Based on the current state s t, this method will select an action to execute. If the current state is state_index, this method selects a node from the nodes connected to the current node according to the probability sequence and adds it to the current path. If the current state is state_create, this method selects to create a new path and selects one of all the links with telemetry requirements as the starting point of the new path.
[0208] Reward: In the algorithm, the reward mechanism is a key factor guiding the model to make optimal decisions. The goal is to minimize the weighted telemetry benefit, that is, it is hoped that the dynamic probe path can meet the most telemetry requirements with the least resource consumption. The reward function calculates a reward value according to the satisfaction of each telemetry requirement under the current path deployment plan. This reward value is composed of the weighted sum of multiple telemetry metrics (such as latency, bandwidth utilization, etc.), and each metric has a corresponding weight, which can be adjusted according to the actual needs of users. The algorithm simulates different path deployment plans and calculates the weighted telemetry benefit for each plan. Then, the plan with the minimum weighted telemetry benefit is selected as the optimal solution and the corresponding reward is given. This reward value is fed back to the model to update the model's parameters so that a better choice can be made in the next decision. The dynamic probe can have multiple possible paths to collect network status information. Each path deployment plan corresponds to different resource consumption and telemetry performance. This algorithm simulates these different path deployment plans and selects the optimal path deployment plan by evaluating the network performance and resource consumption under each plan. The weighted telemetry benefit is a comprehensive metric used to evaluate the overall performance of each path deployment plan. It considers multiple performance metrics, such as control overhead, latency, bandwidth utilization, etc., and each metric has a corresponding weight, reflecting the importance of the metric in the overall evaluation, to find the optimal solution.
[0209] The periodic calculation model aims to model data such as network performance metrics and network topology to predict the sending period of the dynamic probe.
[0210] Model structure:
[0211] Input layer: Receives time series data, including network performance metrics (latency, bandwidth utilization, packet loss rate, etc.), targets, and network topology.
[0212] LSTM layer: One layer of LSTM cells is used to capture long-term dependencies in the time series. Each layer of LSTM contains multiple cells, and dropout is set to prevent overfitting.
[0213] Dense layer: After the LSTM layer, a Dense layer is added to map the output of the LSTM layer to the final predicted values: the number of transmissions and the period.
[0214] Output layer: Output the predicted number of transmissions and intervals.
[0215] This application uses a deep reinforcement learning (DRL) framework to train a reinforcement learning model, and its training process is as follows:
[0216] Initialize the global behavior network Actor Network and the global evaluation network Critic Network: Initialize the global Actor Network and the global Critic Network with random weights. The global Actor Network is responsible for predicting the probability distribution of the next action. The global Critic Network is responsible for estimating the value function in a given state.
[0217] Create multiple task workers: Each worker has its own Actor Network and Critic Network, which are copies of the global network. Each agent runs in an independent training environment.
[0218] Training process of the worker (asynchronous):
[0219] Sample network topology instances: Sample network topologies from the training set, including specific network topologies and related node information.
[0220] Generate paths: Generate possible dynamic probe paths according to the current policy (output by the agent's Actor Network). The generation of paths is a sequential decision-making process, and the next node is selected at each step based on the current state.
[0221] Calculate rewards: Calculate rewards based on the generated paths and the user's optimization objectives. The reward function is calculated according to the performance metrics and weights of the paths.
[0222] When the value of the reward function increases, the system updates the bitmap according to the new optimization objective, reducing the number of transmissions and frequencies to ensure that DPs can collect more data related to key performance indicators.
[0223] When the value of the reward function decreases, the model explores different actions and paths more actively, updates more metrics to be covered, and increases the number of transmissions and frequencies.
[0224] Calculate gradients and update the local network: Use the policy gradient algorithm to calculate the gradients of the Actor Network. Use the Temporal Difference Error to calculate the gradients of the Critic Network. Update the parameters of the agent's Actor Network and Critic Network through backpropagation.
[0225] Push gradients to the global network: The agent pushes the gradients it calculates to the global network. The global network uses these gradients to update its own parameters.
[0226] Pull the latest parameters from the global network: The worker pulls the latest parameters from the global network to update its own Actor network and Critic network.
[0227] Repeat the above steps: All agents asynchronously repeat the above steps until the model converges.
[0228] Model convergence: When the model converges, use the global Actor network and the global Critic network as the final model.
[0229] The loss function of the reinforcement learning model is as follows:
[0230] ;
[0231] In the formula, N is the number of satellite nodes included in the path, L s represents the resource information occupied by the probe packet corresponding to node s, P s is the frequency corresponding to the satellite node; n is the number of performance indicators included in the network status information, w i represents the weight value corresponding to the i-th performance indicator, m i represents the i-th performance indicator value.
[0232] Taking the conditional probability part training explanation as an example, it is explained as follows:
[0233] Assume that the current satellite node is, and it is connected to three network satellite nodes A, B, and C. The network status information of these three is included in X t in, X t is the status information of the current node, including its connection status with adjacent nodes, link delay information, etc.
[0234] The n in the formula refers to the number of samplings in the conditional probability model training, which is set manually.
[0235] Conditional probability formula part:
[0236] .
[0237] u t is the unnormalized score, where w c is the trainable weight matrix, which linearly transforms the concatenation result of x and h; v c is the trainable variable used to calculate the unnormalized score.
[0238] , c t is used to calculate the hidden state. θ is the updated parameter in the actor network. When updating θ, w c and vc Update together.
[0239] The Actor network is used to calculate the conditional probability, and the Critic network is used to evaluate the reward.
[0240] Initialize the global Actor network parameters θ and Critic network parameters φ (both parameters are randomly generated). Initialize the global gradients △θ = 0, △φ = 0.
[0241] Parallel training: Create multiple workers, each with its own Actor network and Critic network, which are copies of the global network. Each worker runs in an independent training environment.
[0242] The training process for each worker:
[0243] Initialization: Each worker fetches the latest parameters θ and φ from the global network.
[0244] Sampling: Sample in the current environment. For each sampling, select a path through the conditional probability and obtain a new Based on the selected path, calculate the reward .
[0245] Update the gradient:
[0246] ;
[0247] Synchronize the gradients to the global network: Synchronize the calculated gradients △θ and △φ to the global network.
[0248] Update the global network parameters:
[0249] ;
[0250] Update the network parameters of the worker: Fetch the latest parameters θ and φ from the global network and update the worker's network.
[0251] Termination condition: Stop training when the maximum number of training rounds is reached or the performance of the global network no longer improves.
[0252] When the electronic device collects probe data, the telemetry analyzer parses each probe into a series of dictionaries or tuples. The dictionary or tuple contains specific telemetry information collected from various nodes in the network (such as switches, satellites, etc.), such as link latency, queue length, port throughput, etc. It can be used for network monitoring, fault diagnosis, and traffic control. For basic measurement data, the analyzer provides the basic network status for the dynamic probe path system. For the data plane (DP), the analyzer stores the network information in a database, which can be accessed by other network applications. To ensure structured storage, the storage format is designed according to the switch ID and metadata type. To save storage space, whenever new probe data is received, the telemetry analyzer updates the content of the existing dictionary or tuple with the latest network information. Finally, the telemetry analyzer calculates the high-level metadata based on the query. Network applications can obtain network telemetry reports regularly according to their telemetry requirements. High-level metadata refers to higher-level network performance metrics or status information derived from the raw data collected from network nodes. The high-level metadata of link quality is obtained by calculating the average latency of the link over a period of time.
[0253] Another task of telemetry analysis is to locate network faults. When a link fails, all dynamic probes passing through the faulty link cannot be collected by the controller. To minimize the number of dynamic probes that need to be checked, it is necessary to ensure that only one dynamic probe passes through each link. By using non-overlapping dynamic probe paths, the path where the fault lies can be quickly found, and the workload of fault location can be reduced. When generating the shortest geodesic, it can be ensured that a unique probe path is assigned to each link.
[0254] This application proposes an adaptive in-band network telemetry method assisted by dual-time-scale probes, including basic probes and dynamic probes. Technically speaking, the basic probes are set to a long period to collect basic network status information, and the dynamic probes are set to a dynamic period to adaptively and specifically collect network status information according to the telemetry task input and network status.
[0255] In this application, the dynamic in-band measurement module adopts the Actor-Critic mechanism based on deep reinforcement learning. It can optimize the future path planning strategy according to the previously collected network information and path planning results, and can automatically complete these tasks, reducing the dependence on manual intervention.
[0256] The telemetry analysis module in this application proposes a fault location mechanism. When a link fails, all dynamic probes passing through the faulty link cannot be collected by the controller. To minimize the number of dynamic probes that need to be checked, it is necessary to ensure that only one dynamic probe passes through each link. By using non-overlapping dynamic probe paths, the path where the fault lies can be quickly found, and the workload of fault location can be reduced.
[0257] This application proposes an adaptive in-band network telemetry method assisted by dual time-scale probes. By using basic probes to collect basic network status information and dynamic probes to adaptively and specifically collect network status information based on the input of telemetry tasks and the results of self-learning of network status, it can effectively improve the adaptability and flexibility of in-band measurement, ensure the in-band measurement effect, and reduce the in-band measurement overhead.
[0258] In this application, the dynamic in-band measurement module adopts the Actor-Critic mechanism based on deep reinforcement learning. It can automatically adjust the deployment plan according to different requirements of users, is more robust than traditional heuristic algorithms, enables the telemetry system to be adaptable, and is no longer restricted by the experience accumulation and subjective judgment of measurement personnel. It can also quickly solve the same type of combinatorial optimization problems, adapt to the frequent fluctuations of network load and dynamic changes in network environments such as links that may be disconnected or congested, timely adjust system parameters, and optimize the quality and efficiency of network telemetry. At the same time, its self-learning ability can effectively extract key features in the collected data, guide the analysis results to be free from the influence of noise and interference, and improve the accuracy of measurement and analysis.
[0259] In this application, the telemetry analysis module proposes a fault location mechanism that makes full use of the self-learning mechanism of the dynamic in-band measurement module, fully explores the deep information of historical measurement data, and dynamically guides the discovery and location of network faults. By using non-overlapping dynamic probe paths, it can quickly find the path where the fault is located and reduce the workload and measurement overhead of fault location.
[0260] The measurement method proposed in this application can collect data at the node level, port level, and bitmap level: at the node level, it can measure signal strength, link quality, congestion situation, etc. At the port level, it can measure data transmission rate, bandwidth utilization rate, bit error rate, packet loss rate, delay, and jitter. The bitmap level focuses on signal strength, signal-to-noise ratio, link quality, and queue length. This method can fill the demand gap that traditional telemetry methods cannot meet the index requirements of satellite communication networks, and at the same time, adaptively adjust measurement indicators according to measurement requirements.
[0261] When the telemetry requirements change or the network environment changes significantly (such as changes in network scale), the algorithm proposed by the dynamic in-band measurement module can adopt the transfer learning method, that is, set different reward functions and training sets to assist in training the model to solve different problems and apply it to various network environments.
[0262] Figure 9 The following is a schematic structural diagram of the in-band measurement device provided by this application, including:
[0263] The first acquisition module 11 is used to send a first probe message to at least one first satellite and acquire the first network status information of the at least one first satellite;
[0264] A determination module 12, configured to generate second probe messages corresponding to at least one second satellite based on first network status information of the at least one first satellite.
[0265] A second acquisition module 13, configured to send corresponding second probe messages to the at least one second satellite and acquire second network status information of the at least one second satellite.
[0266] The determination module 12 is specifically configured to input the first network status information of the at least one first satellite into a reinforcement learning model, and generate second probe messages corresponding to at least one second satellite based on the reinforcement learning model.
[0267] The determination module 12 is further configured to input the first network status information of the at least one first satellite into a reinforcement learning model, and determine at least one detection path based on the reinforcement learning model.
[0268] The determination module 12 is specifically configured to send corresponding second probe messages to at least one second satellite in the at least one detection path.
[0269] The first acquisition module 11 is specifically configured to send first probe messages to the at least one first satellite based on a first frequency.
[0270] The second acquisition module 13 is specifically configured to send corresponding second probe messages to the at least one second satellite based on a second frequency.
[0271] Wherein, the second frequency is less than the first frequency.
[0272] The determination module 12 is further configured to input the first network status information of the at least one first satellite into a reinforcement learning model, and determine the second frequency corresponding to the at least one second satellite based on the reinforcement learning model.
[0273] The determination module 12 is further configured to determine bitmap information corresponding to the network status information of the at least one second satellite based on the reinforcement learning model.
[0274] The second acquisition module 13 is specifically configured to send corresponding second probe messages carrying the bitmap information to the at least one second satellite; wherein, the bitmap information is used to indicate measured network status information.
[0275] The first acquisition module 11 is specifically configured to, for the at least one first satellite, determine a target gateway station closest to the first satellite; and send the first probe message to the first satellite through the target gateway station.
[0276] The first acquisition module 11 is specifically configured to obtain, for the at least one first satellite, first network state information of at least one first satellite on the path from the target gateway station to the first satellite.
[0277] The second acquisition module 13 is specifically configured to determine, for the at least one detection path, second network state information of at least one second satellite on the detection path; wherein, the second frequencies corresponding to the second satellites on the detection path are the same.
[0278] The reinforcement learning model is used to determine the next-hop satellite of the satellite, and the at least one detection path includes the satellite and the next-hop satellite; wherein, the reinforcement learning model is used to determine, for the satellite, the conditional probability of the satellite and at least one candidate satellite for the next-hop connection; and determine the candidate satellite with the maximum conditional probability as the next-hop satellite of the satellite.
[0279] The determination module 12 is specifically configured to determine, according to the first network state information of the satellite, at least one satellite adjacent to the satellite and communicating normally, and determine the at least one satellite as at least one candidate satellite.
[0280] The device further includes:
[0281] The training module 14 is configured to determine, for at least one first satellite in at least one sample path, a first reward value corresponding to the first satellite according to the bitmap information and weight value corresponding to the network state information of the first satellite in the current iterative training process, and the transmission frequency of the probe packet; determine a second reward value corresponding to the sample path according to the first reward values corresponding to at least one first satellite in the sample path; update the gradient of the reinforcement learning model, the bitmap information and weight value corresponding to the network state information, and the transmission frequency of the probe packet according to the minimum second reward value corresponding to the at least one sample path; and repeat the iterative training until the reinforcement learning model is trained.
[0282] The training module 14 is specifically configured to determine a first sub-reward value and the resource information occupied by the probe packet according to the bitmap information and weight value corresponding to the network state information of the first satellite in the current iterative process; determine a second sub-reward value according to the resource information occupied by the probe packet and the transmission frequency of the probe packet; and determine the first reward value corresponding to the first satellite according to the first sub-reward value and the second sub-reward value.
[0283] The training module 14 is further configured to determine corresponding associated network state information according to the telemetry requirements of in-band measurement; and determine the initial bitmap information and initial weight value corresponding to the network state information of the first satellite according to the associated network state information.
[0284] The present application also provides an electronic device, such as Figure 10 shown, which includes: a processor 21, a communication interface 22, a memory 23, and a communication bus 24. Among them, the processor 21, the communication interface 22, and the memory 23 communicate with each other through the communication bus 24;
[0285] The memory 23 stores a computer program. When the program is executed by the processor 21, the processor 21 is caused to execute at least one of the above method steps.
[0286] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0287] The communication interface 22 is used for communication between the above electronic device and other devices.
[0288] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0289] The above processor may be a general-purpose processor, including a central processor, a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an application-specific integrated circuit, a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0290] The present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program executable by an electronic device. When the program runs on the electronic device, the electronic device is caused to execute at least one of the above method steps.
[0291] The present application provides a computer program product. The computer program product includes an executable program, and when the executable program is executed by a processor, the above-mentioned method is implemented.
[0292] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn of the basic creative concept. Therefore, the appended claims are intended to be interpreted to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.
[0293] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. An in-band measurement method, characterized in that, Applied to an electronic device, the method includes: Sending a first probe message to at least one first satellite and obtaining first network status information of the at least one first satellite; Inputting the first network status information of the at least one first satellite into a reinforcement learning model, and generating a second probe message corresponding to at least one second satellite based on the reinforcement learning model; Sending the corresponding second probe message to the at least one second satellite and obtaining second network status information of the at least one second satellite; The training process of the reinforcement learning model includes: For at least one first satellite in at least one sample path, determining a first reward value corresponding to the first satellite according to the bitmap information and weight value corresponding to the network status information of the first satellite in the current iterative training process, and the sending frequency of the probe message; determining a second reward value corresponding to the sample path according to the first reward values corresponding to at least one first satellite in the sample path; updating the gradient of the reinforcement learning model, the bitmap information and weight value corresponding to the network status information, and the sending frequency of the probe message according to the minimum second reward value corresponding to the at least one sample path; repeating the iterative training until the reinforcement learning model is trained.
2. The method according to claim 1, wherein The method further includes: Inputting the first network status information of the at least one first satellite into a reinforcement learning model and determining at least one detection path based on the reinforcement learning model; The method includes: Sending the corresponding second probe message to at least one second satellite in the at least one detection path.
3. The method according to claim 2, characterized in that, The method includes: Sending a first probe message to the at least one first satellite based on a first frequency; Sending the corresponding second probe message to the at least one second satellite based on a second frequency; Wherein, the second frequency is less than the first frequency.
4. The method according to claim 3, wherein The method further includes: Inputting the first network status information of the at least one first satellite into a reinforcement learning model and determining the second frequency corresponding to the at least one second satellite based on the reinforcement learning model.
5. The method according to claim 1, wherein The method includes: Determining bitmap information corresponding to the network status information of the at least one second satellite based on the reinforcement learning model; Sending the corresponding second probe message carrying the bitmap information to the at least one second satellite; wherein, the bitmap information is used to indicate the measured network status information.
6. The method according to claim 1, wherein The sending a first probe message to at least one first satellite includes: For the at least one first satellite, determining a target gateway station closest to the first satellite; Sending the first probe message to the first satellite through the target gateway station.
7. The method according to claim 6, wherein Obtaining first network status information of the at least one first satellite includes: For the at least one first satellite, obtaining first network status information of at least one first satellite on the path from the target gateway station to the first satellite.
8. The method according to claim 3, wherein Obtaining second network status information of the at least one second satellite includes: For the at least one detection path, determining second network status information of at least one second satellite on the detection path; wherein, the second frequencies corresponding to the second satellites on the detection path are the same.
9. The method according to claim 2, wherein Determining at least one detection path based on the reinforcement learning model includes: The reinforcement learning model is used to determine the next-hop satellite of the satellite, and the at least one detection path includes the satellite and the next-hop satellite; wherein, the reinforcement learning model is used to determine, for the satellite, the conditional probability of the satellite and at least one candidate satellite for the next-hop connection; and the candidate satellite with the maximum conditional probability is determined as the next-hop satellite of the satellite.
10. The method according to claim 9, wherein The process of determining the candidate satellites for the at least one next-hop connection includes: According to the first network state information of the satellite, determine at least one satellite adjacent to the satellite and in normal communication, and determine the at least one satellite as at least one candidate satellite.
11. The method according to claim 1, characterized in that, The determining of the first reward value corresponding to the first satellite according to the bitmap information and weight value corresponding to the network state information of the first satellite in the current iteration process, and the sending frequency of the probe packet includes: According to the bitmap information and weight value corresponding to the network state information of the first satellite in the current iteration process, determine the first sub-reward value and the resource information occupied by the probe packet. According to the resource information occupied by the probe packet and the sending frequency of the probe packet, determine the second sub-reward value. According to the first sub-reward value and the second sub-reward value, determine the first reward value corresponding to the first satellite.
12. The method according to claim 1, wherein, The method further includes: Determine the corresponding associated network state information according to the telemetry requirements of in-band measurement. According to the associated network state information, determine the initial bitmap information and initial weight value corresponding to the network state information of the first satellite.
13. An in-band measurement device, characterized in that, Applied to an electronic device, the apparatus includes: A first acquisition module, configured to send a first probe packet to at least one first satellite and acquire the first network state information of the at least one first satellite. A determination module, configured to input the first network state information of the at least one first satellite into the reinforcement learning model, and generate a second probe packet corresponding to at least one second satellite based on the reinforcement learning model. A second acquisition module, configured to send the corresponding second probe packet to the at least one second satellite and acquire the second network state information of the at least one second satellite. The apparatus further includes: A training module, configured to, for at least one first satellite in at least one sample path, determine the first reward value corresponding to the first satellite according to the bitmap information and weight value corresponding to the network state information of the first satellite in the current iterative training process, and the sending frequency of the probe packet; determine the second reward value corresponding to the sample path according to the first reward values corresponding to the at least one first satellite in the sample path; update the gradient of the reinforcement learning model, the bitmap information and weight value corresponding to the network state information, and the sending frequency of the probe packet according to the minimum second reward value corresponding to the at least one sample path; and repeat the iterative training until the reinforcement learning model is trained.
14. The device according to claim 13, characterized in that, The determination module is further configured to input the first network state information of the at least one first satellite into the reinforcement learning model and determine at least one detection path based on the reinforcement learning model. The determining module is specifically configured to send corresponding second probe messages to at least one second satellite in the at least one detection path.
15. The device according to claim 14, wherein, The first obtaining module is specifically configured to send first probe messages to the at least one first satellite based on a first frequency; The second obtaining module is specifically configured to send corresponding second probe messages to the at least one second satellite based on a second frequency; Wherein, the second frequency is less than the first frequency.
16. The device according to claim 15, characterized in that, The determining module is further configured to input the first network state information of the at least one first satellite into a reinforcement learning model, and determine the second frequency corresponding to the at least one second satellite based on the reinforcement learning model.
17. The device according to claim 13, characterized in that, The determining module is further configured to determine bitmap information corresponding to the network state information of the at least one second satellite based on the reinforcement learning model; The second obtaining module is specifically configured to send the corresponding second probe messages carrying the bitmap information to the at least one second satellite; wherein, the bitmap information is used to indicate the measured network state information.
18. The device according to claim 13, characterized in that, The first obtaining module is specifically configured to, for the at least one first satellite, determine a target gateway station closest to the first satellite; and send the first probe message to the first satellite through the target gateway station.
19. The device according to claim 18, characterized in that, The first obtaining module is specifically configured to, for the at least one first satellite, obtain the first network state information of at least one first satellite on the path from the target gateway station to the first satellite.
20. The device according to claim 15, characterized in that, The second obtaining module is specifically configured to, for the at least one detection path, determine the second network state information of at least one second satellite on the detection path; wherein, the second frequencies corresponding to the second satellites on the detection path are the same.
21. The apparatus according to claim 14, wherein The reinforcement learning model is used to determine the next-hop satellite of a satellite. The at least one detection path includes the satellite and the next-hop satellite; wherein, the reinforcement learning model is used to determine, for the satellite, the conditional probability of the satellite and at least one candidate satellite for the next-hop connection, and determine the candidate satellite with the maximum conditional probability as the next-hop satellite of the satellite.
22. The device according to claim 21, characterized in that, The determining module is specifically configured to determine, according to the first network state information of the satellite, at least one satellite adjacent to the satellite and in normal communication, and determine the at least one satellite as at least one candidate satellite.
23. The device according to claim 13, wherein, The training module is specifically configured to determine a first sub-reward value and the resource information occupied by the probe message according to the bitmap information and the weight value corresponding to the network state information of the first satellite in the current iteration process; Determine a second sub-reward value according to the resource information occupied by the probe message and the sending frequency of the probe message; Determine a first reward value corresponding to the first satellite according to the first sub-reward value and the second sub-reward value.
24. The device according to claim 13, characterized in that The training module is further configured to determine corresponding associated network state information according to the telemetry requirements of in-band measurement; and determine the initial bitmap information and the initial weight value corresponding to the network state information of the first satellite according to the associated network state information.
25. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; A memory for storing a computer program; A processor for implementing the method according to any one of claims 1-12 when executing the program stored in the memory.
26. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the method according to any one of claims 1-12 is implemented when the computer program is executed by a processor.
27. A computer program product, characterized in that, The computer program product includes an executable program, and the method according to any one of claims 1 to 12 is implemented when the executable program is executed by a processor.
Citation Information
Patent Citations
Method and device for selecting perception information reporting node of satellite network and storage medium
CN118900149A
Traffic signal data automatic classification control method based on target identification
CN119672993A