A method for adaptive matching of fiber interfaces based on reinforcement learning

CN122601076APending Publication Date: 2026-08-18GANSU ELECTRIC POWER TIANSHUI POWER SUPPLY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610742483.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-27
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

当光纤外径、包覆状态、业务光强弱或夹持位置发生变化时,相同夹持动作产生的漏光响应并不一致,容易出现漏光信号过弱、放大电路饱和、光方向判定不稳定或光功率估算波动较大的问题

Benefits of technology

(1)将夹持动作和电路切换后产生的过渡观测从状态输入中滤除,仅以稳定观测形成初始结算状态和后续结算状态,避免将瞬态波动误判为有效漏光响应,提高了光方向判定和光功率估算的稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122601076A_ABST
    Figure CN122601076A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on reinforcement learning's optical fiber interface adaptive matching method, comprising the following steps: collecting and marking macro-bending light leakage interaction data, generates training sample set;Dreamer world model is constructed, and execution deviation, observation lag and failure trajectory are written into state evolution, and target model is obtained;Collect initial light leakage observation of optical fiber to be measured;Initial settlement state is generated after lag settlement;Candidate command action is generated by target model and is converted into actual equivalent action;Based on actual equivalent action, candidate path is deduced and pruned, and target adjustment action is determined;After execution, re-collect light leakage observation, generate subsequent settlement state and output matching result.The application realizes the adaptive optimization of non-invasive service light detection parameters, improves the stability of optical fiber interface matching, detection accuracy and weak light determination reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of optical fiber detection technology, and in particular to an adaptive matching method for optical fiber interfaces based on reinforcement learning. Background Technology

[0002] In the process of fiber optic operation and maintenance, weak light remediation and interface verification, it is usually necessary to determine whether there is service light in the fiber, the direction of service light transmission and the magnitude of optical power without cutting or pulling out the fiber.

[0003] Existing detection methods typically rely on preset clamping stroke, fixed probe position, and fixed amplification gain. When the fiber outer diameter, cladding state, service light intensity, or clamping position changes, the leakage light response produced by the same clamping action is not consistent, which can easily lead to problems such as weak leakage light signal, amplifier circuit saturation, unstable optical direction determination, or large fluctuations in optical power estimation. Simply increasing the clamping degree to enhance the leakage light signal may also cause excessive fiber bending, increasing the additional loss of the service light.

[0004] Existing methods typically use immediate sampling results directly as the basis for detection after the action is executed, without distinguishing between transient and stable observations caused by clamping actions and circuit switching. This easily leads to the misinterpretation of transient fluctuations as valid light leakage responses. Existing model-based parameter tuning methods also often directly extrapolate the state based on the command action, without considering the execution deviation between the command action and the actual equivalent action. Furthermore, they lack a mechanism for early truncation of failed adjustment paths, resulting in repeated invalid adjustments and making it difficult to quickly obtain stable detection parameters adapted to the current fiber optic interface.

[0005] Therefore, how to provide an adaptive matching method for fiber optic interfaces based on reinforcement learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose an adaptive matching method for fiber optic interfaces based on reinforcement learning. This invention utilizes fiber macrobend leakage detection and the Dreamer world model to adaptively adjust the clamping, probe receiving, amplification gain, and frequency identification parameters of the detection device. It has the advantages of reducing transient misjudgments, reducing ineffective adjustments, and improving the stability of optical direction and optical power detection.

[0007] An adaptive matching method for fiber optic interfaces based on reinforcement learning according to an embodiment of the present invention includes the following steps: Collect fiber macrobend leakage light interaction data, annotate the fiber macrobend leakage light interaction data, and generate a training sample set characterizing the fiber interface detection environment. A Dreamer world model is constructed based on the training sample set. Action execution bias, post-action observation lag and failure adjustment trajectory are written into the state evolution process to train the target Dreamer world model. The fiber under test is placed into the detection slot and an initial clamping action is applied. The fiber under test forms a macro-bend leakage state in the detection area, and the initial leakage observation is collected. The initial light leakage observations are settled with lag. Transient observations caused by clamping actions and circuit switching are identified. Transient observations are filtered out to generate stable observations. The initial settlement state is formed by the stable observations. The initial settlement state is input into the target Dreamer world model to generate candidate command actions. The actual equivalent actions corresponding to the candidate command actions are converted by executing the deviation perception process. The latent space inference is based on the actual equivalent actions. Based on the actual equivalent action, candidate adjustment paths are deduced. Candidate adjustment paths that meet the preset failure criteria are identified through the failure trajectory pruning process. The identified candidate adjustment paths are truncated, and the target adjustment action is determined from the remaining candidate adjustment paths. After performing the target adjustment action, the light leakage observation is re-acquired, and the subsequent settlement status is generated according to the lag settlement. When the subsequent settlement status meets the preset matching criteria, the fiber optic interface matching result is output.

[0008] Optionally, the annotation process for the fiber macrobend leakage data specifically includes: The training fiber is made to form a macro-bend and light leakage state in the detection slot according to the preset clamping action, and the detection action data and light leakage response data are recorded. Baseline correction, outlier sampling removal, and time alignment are performed on the light leakage response data to obtain the effective light leakage sequence corresponding to the detection action data. Based on the response relationship between the detected action data and the effective light leakage sequence, an action execution deviation label is generated; Based on the magnitude and duration of the changes in the effective light leakage sequence after the detection action, observation settlement labels are generated; Based on whether the detection results corresponding to the valid light leakage sequence meet the preset matching criteria, a matching result label is generated; The detection action data, effective light leakage sequence, action execution deviation annotation, observation settlement annotation, and matching result annotation are combined into sample records in chronological order, and a training sample set representing the fiber optic interface detection environment is generated from the sample records.

[0009] Optionally, the step of writing the action execution deviation, post-action observation lag, and failure adjustment trajectory into the state evolution process specifically includes: Read sample records from the training sample set and encode the valid light leakage sequences in the sample records as the current potential state; The detected action data is combined with the action execution deviation annotation to generate the actual equivalent action, which is then used as the action input for the state evolution process. Based on the observation settlement label, stable observations are extracted from the effective light leakage sequence, and stable observations are encoded as the next potential state. Transitional observations are excluded from the next state in the state evolution process. Based on the matching result annotation and preset failure criteria, a failure adjustment trajectory identifier is generated, and the failure adjustment trajectory identifier is written into the trajectory termination condition of the state evolution process; when generating the failure adjustment trajectory identifier, the matching result annotation in the sample record is read first. The Dreamer world model is trained using the current potential state, the actual equivalent action, the next potential state, and the trajectory termination condition to obtain the target Dreamer world model.

[0010] Optionally, the initial light leakage observation specifically includes: Obtain the clamping start signal after the optical fiber under test is placed into the detection slot, and call the initial action command corresponding to the initial clamping action; An initial sampling task is generated based on the initial action command, and the clamping start time, clamping completion time, and sampling start time are recorded in the initial sampling task; Read the raw light leakage response data of the optical fiber under test after the initial clamping action from the output end of the PD photodiode or TO photodetector. The original light leakage response data is truncated according to the sampling start time to obtain the initial sampling segment; Background baseline subtraction and sample value normalization are performed on the initial sampling segment to generate initial light leakage observations; The initial light leakage observation is linked to the initial action command, the clamping start time, and the clamping completion time to form an initial observation record.

[0011] Optionally, the lag settlement of the initial light leakage observation specifically includes: Read the initial light leakage observation, initial action command, clamping start time, and clamping completion time from the initial observation record; The settlement starting point for the initial light leakage observation is taken as the moment the clamping is completed. The sampling content generated by the initial clamping action before the settlement starting point is marked as transitional observation. After the settlement start point, the initial light leakage observations are continuously read in chronological order. Combined with the circuit switching process corresponding to the initial action command, the short-term sudden sampling content caused by the circuit switching is identified and incorporated into the transition observation. The continuity of the remaining initial light leakage observations is judged. When the continuously read sampling content maintains a stable change state relative to the previous sampling content, the corresponding sampling content is determined as a stable observation. If the continuously read sampled content still shows clamping recovery fluctuations or circuit switching fluctuations, it will continue to be marked as a transitional observation and the settlement will be delayed. Transient observations are filtered out, and only stable observations are retained as the state input source for the target Dreamer world model. The initial settlement state is generated based on the stable observations.

[0012] Optionally, the latent space deduction based on actual equivalent actions specifically includes: The initial settlement state is input into the target Dreamer world model, which then generates the first round of candidate command actions based on the initial settlement state. Based on the action execution deviation in the state evolution process, the execution amount of the first round of candidate command actions is corrected to obtain the first round of actual equivalent actions. The first round of candidate command actions are replaced by the first round of actual equivalent actions as the action input for the state evolution process. The target Dreamer world model starts from the initial settlement state and predicts the first predicted settlement state after the first round of actual equivalent actions. Using the first predicted settlement state as the starting point for the second round of deduction, the target Dreamer world model generates the second round of candidate command actions. The execution amount is then corrected again based on the action execution deviation to obtain the second round of actual equivalent actions. The second round of actual equivalent actions are used as the action input for the state evolution process to predict the second predicted settlement state. Candidate command actions, actual equivalent actions, and predicted settlement states are continuously generated in the above manner to form a candidate adjustment path that is sequentially connected by several predicted settlement states; In each round of potential space simulation, candidate command actions do not directly participate in the state evolution process; the state evolution process is predicted using the actual equivalent actions of the corresponding round. During the application phase, the model parameters of the target Dreamer world model remain unchanged, and the deviation perception process is only used to correct the action inputs used in the latent space extrapolation.

[0013] Optionally, determining the target adjustment action from the remaining candidate adjustment paths specifically includes: Based on the actual equivalent actions, several candidate adjustment paths are generated in the latent space of the target Dreamer world model. Each candidate adjustment path consists of the predicted settlement state and the actual equivalent action connected in the deduction order. The predicted settlement status is read in the order of deduction for each candidate adjustment path, and the predicted settlement status is matched with the preset failure criteria to generate the failure identification result for each candidate adjustment path. When any predicted settlement state in the candidate adjustment path meets the preset failure criterion, the corresponding candidate adjustment path will be truncated from the predicted settlement state that meets the preset failure criterion, and subsequent predicted settlement states after the truncation will no longer be generated. Candidate adjustment paths that are to be truncated are excluded from the selection range of target adjustment actions, and the starting settlement state, actual equivalent action sequence and failure identification result corresponding to the truncated candidate adjustment path are combined into a pruning template; When the candidate adjustment path generated by subsequent potential space deduction has the same initial settlement state range and the same actual equivalent action sequence as the pruning template, stop deducing the subsequent predicted settlement state. The candidate adjustment path that first satisfies the preset matching criterion is selected from the candidate adjustment paths that have never been truncated. The candidate command action corresponding to the first actual equivalent action of the selected candidate adjustment path is determined as the target adjustment action.

[0014] Optionally, the step of outputting the fiber optic interface matching result when the preset matching criteria are met in the subsequent settlement state specifically includes: After the target adjustment action is performed, leakage observations are collected again, and the completion time of the target adjustment action is used as the settlement starting point for subsequent lag settlements. According to the delayed settlement method, the leaky light observations that are re-acquired are identified as transitional observations and stable observations are determined. The sampling content that is still affected by the target adjustment action and circuit switching process after the settlement starting point is marked as transitional observations. After the transitional observation, when the continuously read leaky light observations enter a stable change state, the corresponding leaky light observations are determined as subsequent stable observations; Filter out transitional observations and encode subsequent stable observations in the state input format of the target Dreamer world model to generate subsequent settlement states; The subsequent settlement status is compared with the preset matching criteria. If the subsequent settlement status meets the preset matching criteria, the fiber optic interface matching result is output. If the subsequent settlement status does not meet the preset matching criteria, the subsequent settlement status is used as the input status for the next round.

[0015] The beneficial effects of this invention are: (1) The transient observations generated after clamping action and circuit switching are filtered out from the state input, and only the stable observations are used to form the initial settlement state and subsequent settlement state, so as to avoid misjudging transient fluctuations as effective light leakage response and improve the stability of light direction determination and light power estimation.

[0016] (2) The candidate command actions generated by the target Dreamer world model are converted into actual equivalent actions, and the potential space is extrapolated using actual equivalent actions. This avoids the state prediction deviation caused by the inconsistency between the command actions and the actual response of the equipment, and improves the accuracy of adaptive adjustment of detection parameters.

[0017] (3) When the candidate adjustment path meets the preset failure criteria, the corresponding path is cut off in advance, and the target adjustment action is determined from the remaining candidate adjustment paths. This reduces invalid adjustment processes such as saturation, directional instability, and frequency loss, and improves the matching efficiency and detection reliability of the fiber optic interface. Attached Figure Description

[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of an adaptive matching method for fiber optic interfaces based on reinforcement learning proposed in this invention. Figure 2 This is a schematic diagram illustrating the delayed settlement generation of settlement state in a fiber optic interface adaptive matching method based on reinforcement learning proposed in this invention. Figure 3 This is a schematic diagram illustrating the execution deviation perception deduction of the fiber optic interface adaptive matching method based on reinforcement learning proposed in this invention. Detailed Implementation

[0019] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0020] refer to Figures 1-3 An adaptive matching method for fiber optic interfaces based on reinforcement learning includes the following steps: Collect fiber macrobend leakage light interaction data, annotate the fiber macrobend leakage light interaction data, and generate a training sample set characterizing the fiber interface detection environment. A Dreamer world model is constructed based on the training sample set. Action execution bias, post-action observation lag and failure adjustment trajectory are written into the state evolution process to train the target Dreamer world model. The fiber under test is placed into the detection slot and an initial clamping action is applied. The fiber under test forms a macro-bend leakage state in the detection area, and the initial leakage observation is collected. The initial light leakage observations are settled with lag. Transient observations caused by clamping actions and circuit switching are identified. Transient observations are filtered out to generate stable observations. The initial settlement state is formed by the stable observations. The initial settlement state is input into the target Dreamer world model to generate candidate command actions. The actual equivalent actions corresponding to the candidate command actions are converted by executing the deviation perception process. The latent space inference is based on the actual equivalent actions. Based on the actual equivalent action, candidate adjustment paths are deduced. Candidate adjustment paths that meet the preset failure criteria are identified through the failure trajectory pruning process. The identified candidate adjustment paths are truncated, and the target adjustment action is determined from the remaining candidate adjustment paths. After performing the target adjustment action, the light leakage observation is re-acquired, and the subsequent settlement status is generated according to the lag settlement. When the subsequent settlement status meets the preset matching criteria, the fiber optic interface matching result is output.

[0021] In this embodiment, the annotation processing of optical fiber macrobend leakage data specifically includes: The training optical fiber is subjected to a preset clamping action to form a macro-bend with light leakage within the detection slot, and the detection action data and light leakage response data are recorded. The preset clamping action can be performed using a non-invasive fiber clamping method, in which the training optical fiber is inserted into the detection slot or cable slot, and then the clamping mechanism automatically clamps the training optical fiber, forming a macro-bend with light leakage without blocking the transmission of service light. After the training optical fiber forms a macrobend, some of the light transmitted inside the training optical fiber leaks along the bend. A PD photodiode or TO photodetector placed on one side of the bend collects the leaked light and converts it into light leakage response data. The detection action data includes the clamping displacement, clamping time, probe receiving angle, amplification circuit setting, and sampling time window corresponding to the clamping action; the light leakage response data includes the electrical signal output by the PD photodiode or TO photodetector, the light leakage intensity calculated from the electrical signal, the light direction determination result, and the optical power estimation result. Baseline correction, outlier removal, and time alignment are performed on the light leakage response data to obtain the effective light leakage sequence corresponding to the detection action data. Background light leakage response is collected when the tablet pressing mechanism is not clamping the training fiber, and the average value of the background light leakage response is used as the baseline value. The baseline value is subtracted from the light leakage response data after the clamping action to obtain the corrected light leakage response data. The corrected light leakage response data is checked for continuity according to the sampling time, and unstable samples caused by single-point abrupt changes or short-term full-scale saturation are removed as outliers. Missing positions after removal are filled in by adjacent effective samples. The time of completion of the clamping action is used as the time reference to align the corrected light leakage response data with the detection action data to form the effective light leakage sequence. Based on the response relationship between the detected action data and the effective light leakage sequence, an action execution deviation label is generated; the command action amplitude in the detected action data is read, and the stable response before and after the action is extracted from the effective light leakage sequence, and the change in light leakage response before and after the action is calculated; according to the pre-obtained correspondence between the command action and the change in light leakage response, the change in light leakage response is converted into the actual equivalent action amount; the difference between the command action amplitude and the actual equivalent action amount is used as the action execution deviation label; the action execution deviation label is used to record the actual response differences generated after execution of tablet clamping, probe receiving adjustment, or amplification gain switching. Based on the change amplitude and duration of the effective light leakage sequence after the detection action, observation settlement labels are generated. Starting from the completion time of the clamping action, continuous window detection is performed on the effective light leakage sequence. When the change amplitude of the light leakage response within multiple consecutive sampling windows remains within a preset stable range, the corresponding sampling segment is labeled as a stable observation. The sampling segment before the appearance of a stable observation is labeled as a transitional observation. If no stable observation appears after a preset waiting time, the corresponding effective light leakage sequence is labeled as an unsettled observation. The observation settlement labels are used to determine the stable observations in the training sample set that can be used as the results of state evolution. Based on whether the detection results corresponding to the valid light leakage sequence meet the preset matching criteria, a matching result label is generated; the light leakage intensity, signal-to-noise ratio, light direction determination result, and light power estimation result are extracted from stable observations; if the light leakage intensity is within the effective detection range, the signal-to-noise ratio meets the detection requirements, the light direction determination result is continuous and stable, the light power estimation result fluctuates within the preset range, and no continuous saturation output is generated during the clamping process, a matching success label is generated; if any condition is not met, or the valid light leakage sequence is labeled as an unsettled observation, a matching failure label is generated; in the fixed frequency detection scenario, a matching success label also requires the presence of the target frequency component in the stable observations; The detection action data, effective light leakage sequence, action execution deviation annotation, observation settlement annotation, and matching result annotation are combined into sample records in chronological order, and a training sample set representing the fiber optic interface detection environment is generated from the sample records.

[0022] In this embodiment, incorporating action execution deviation, post-action observation lag, and failure adjustment trajectory into the state evolution process specifically includes: Sample records are read from the training sample set, and the effective light leakage sequences in the sample records are encoded as the current potential state. The effective light leakage sequence segments are divided into fixed-length sampling windows, each containing 10 consecutive sampling points. The mean light leakage response, the rate of change of light leakage response, the signal-to-noise ratio, and the output saturation flag are calculated for each sampling window. The calculation results are input into the observation encoder of the Dreamer world model to generate the current potential state. The detection action data and action execution deviation annotation are combined to generate the actual equivalent action, which is used as the action input for the state evolution process. The actual equivalent action is calculated by subtracting the command action amplitude from the action data and the action execution deviation annotation. The command action amplitude is the control quantity corresponding to the tablet clamping action, probe receiving adjustment action, or amplification gain switching action. The action execution deviation annotation is obtained from the response relationship between the detection action data and the effective light leakage sequence. The state evolution process does not directly use the command action amplitude, but uses the actual equivalent action as the action input. Stable observations are extracted from the effective leaked light sequence based on the observation settlement labels, and encoded as the next potential state. Transitional observations are excluded from the next state in the state evolution process. The start and end times of stable observations are read according to the observation settlement labels, and the corresponding sampling segments are extracted from the effective leaked light sequence as stable observations. When the observation settlement label is an unsettled observation, no next potential state is generated. The same sampling window processing as the current potential state is performed on the stable observations to obtain stable observation features. The stable observation features are input into the observation encoder to generate the next potential state. The sampling segment between the action completion time and the stable observation start time is used as a transitional observation. Transitional observations are only retained in the sample record and are not used as the next potential state in the state evolution process. A failure adjustment trajectory identifier is generated based on the matching result annotation and preset failure criteria. This identifier is then written into the trajectory termination condition of the state evolution process. When generating the failure adjustment trajectory identifier, the matching result annotation in the sample record is read first. If the matching result annotation is a failure annotation, stable observations, unsettled observations, and valid light leakage sequences in the same sample record are read. The preset failure criteria can be set as follows: ADC output continuously exceeds 95% of full scale; the light direction determination result changes within three consecutive sampling windows; the light leakage response fails to enter a stable range within a preset waiting time; and the optical power estimation result changes by more than 0.2 dB within a consecutive window. If any preset failure criterion is met, the failure adjustment trajectory identifier is set to 1; if the preset failure criterion is not met, the failure adjustment trajectory identifier is set to 0. Sample records with a failure adjustment trajectory identifier of 1 are written into the trajectory termination condition during the state evolution process. The Dreamer world model is trained using the current potential state, the actual equivalent action, the next potential state, and the trajectory termination condition to obtain the target Dreamer world model. During training, the current potential state is used as the starting point of state evolution, the actual equivalent action is used as the action input, the next potential state is used as the state evolution target, and the trajectory termination condition is used as the end marker of the candidate adjustment path. When the trajectory termination condition indicates that the trajectory has not terminated, the Dreamer world model continues to learn the state evolution relationship from the current potential state to the next potential state. When the trajectory termination condition indicates that the trajectory has terminated, the Dreamer world model stops generating subsequent imagined trajectories along the corresponding sample record. After training is completed, the target Dreamer world model is obtained. The target Dreamer world model is used to generate candidate command actions and deduce candidate adjustment paths based on the settlement state during the application phase.

[0023] In this embodiment, the initial light leakage observation specifically includes: The clamping start signal is obtained after the fiber under test is placed into the detection slot, and the initial action command corresponding to the initial clamping action is called. The clamping start signal can be generated by the detection slot position switch, the fiber placement detection signal or the user trigger command. The initial action command is used to make the fiber under test form a non-intrusive macro-bend light leakage state. The initial action command can include the initial clamping stroke and the clamping holding time. An initial sampling task is generated based on the initial action command. The initial sampling task records the clamping start time, clamping completion time, and sampling start time. The initial sampling task is used to synchronously collect light leakage response data during the execution of the initial clamping action. The clamping start time is the time point when the initial action command begins to be executed, the clamping completion time is the time point when the tablet pressing mechanism reaches the initial clamping stroke, and the sampling start time can be set to 50ms before the clamping start time to retain the background response data before the clamping action. The original light leakage response data of the optical fiber under test after the initial clamping action is read from the output end of the PD photodiode or TO photodetector. The original light leakage response data can be the voltage sampling sequence after the output end of the PD photodiode or TO photodetector is converted by the amplifier circuit. The sampling frequency can be set to 10kHz, and the sampling duration can cover the time interval from the sampling start time to 500ms after the clamping completion time, so as to retain the background response before clamping, the transition response of clamping action, and the light leakage response after clamping. The original light leakage response data is truncated according to the sampling start time to obtain the initial sampling segment; the sampling start time is taken as the starting point of the initial sampling segment, and the preset delay time after the clamping completion time is taken as the ending point of the initial sampling segment; the initial sampling segment includes the background sampling part before the clamping action and the light leakage sampling part after the clamping action; Background baseline subtraction and sample value normalization are performed on the initial sampling segment to generate initial light leakage observation. When subtracting the background baseline, the background sampling part before the clamping start time can be selected, the mean value of the background sampling can be calculated as the background baseline value, and the background baseline value can be subtracted from each sample value in the initial sampling segment. When normalizing the sample value, the sample value after subtracting the background baseline can be normalized according to the current gain level of the amplifier circuit and the ADC range, so that the initial sampling segments under different gain levels are mapped to a uniform numerical range. The initial light leakage observation is linked to the initial action command, the clamping start time, and the clamping completion time to form an initial observation record.

[0024] In this embodiment, performing delayed settlement on the initial light leakage observation specifically includes: Read the initial light leakage observation, initial action command, clamping start time, and clamping completion time from the initial observation record; The settlement starting point for the initial light leakage observation is the clamping completion time. The sampling content generated by the initial clamping action before the settlement starting point is marked as transitional observation. The clamping completion time can be the moment when the tablet pressing mechanism reaches the clamping position corresponding to the initial action command. The sampling content before the settlement starting point includes the light leakage response changes from the start of clamping to the completion of clamping. This part of the sampling content is not directly used as the state input of the target Dreamer world model. A minimum waiting time of 20ms can be set after the clamping completion time. The sampling content within the minimum waiting time is still classified as transitional observation to avoid the short mechanical recovery process after tablet clamping. After the settlement start point, the initial light leakage observations are continuously read in chronological order. Combined with the circuit switching process corresponding to the initial action command, the short-term sudden sampling content caused by the circuit switching is identified and incorporated into the transition observation. The circuit switching process is the amplification gain switching, sampling range switching, or filter parameter switching involved in the initial action command. When identifying the short-term sudden sampling content, the initial light leakage observations within 3 sampling windows after the circuit switching occurs are read with the time of circuit switching as the center. If the change amplitude of adjacent sampling values ​​exceeds 20% of the stable average value before switching, or the sampling value shows full-scale output, the corresponding sampling content is determined as the short-term sudden sampling content. For the remaining initial light leakage observations, a continuity judgment is made. When the continuously read sampling content maintains a stable change state relative to the previous sampling content, the corresponding sampling content is determined to be a stable observation. The continuity judgment is performed according to a fixed sampling window. When the change range of the window mean of no less than 3 consecutive sampling windows is less than 5% of the mean of the previous sampling window, and the difference between the maximum and minimum values ​​within the window is less than 5% of the window mean, it is determined that the continuously read sampling content maintains a stable change state, and the sampling content that meets the conditions is determined to be a stable observation. If the continuously read sampling content still exhibits clamping recovery fluctuations or circuit switching fluctuations, it is marked as a transitional observation and the settlement is delayed. Clamping recovery fluctuations can be manifested as the mean of the continuous sampling window exceeding the threshold for an extended period. Circuit switching fluctuations can be manifested as short-term jumps, full-scale output, or inability to form a continuous stable window after the circuit switching occurs. When delaying the settlement, the sampling window is moved forward segment by segment until a stable observation appears. If a stable observation is not formed within 3 seconds after the clamping completion time, the initial light leakage observation is recorded as an unsettled observation. Transient observations are filtered out, and only stable observations are retained as the state input source for the target Dreamer world model. The initial settlement state is generated based on the stable observations.

[0025] In this embodiment, the latent space deduction is based on actual equivalent actions and specifically includes: The initial settlement state is input into the target Dreamer world model, which then generates the first round of candidate command actions based on the initial settlement state. These first round of candidate command actions are generated by the policy network of the target Dreamer world model, and each candidate command action corresponds to one of the following adjustment types: tablet clamping, probe reception, amplification gain, and frequency identification. Five candidate command actions are generated in each round. Based on the action execution deviation in the state evolution process, the execution amount of the first round of candidate command actions is corrected to obtain the first round of actual equivalent actions. The action execution deviation is obtained from the action execution deviation annotation during the training phase. When correcting the execution amount, the command action amplitude of the first round of candidate command actions is read, and the corresponding adjustment type of action execution deviation is subtracted from the command action amplitude to obtain the first round of actual equivalent actions. If the first round of candidate command actions is a tablet clamping action, the correction object is the tablet displacement; if the first round of candidate command actions is a probe receiving action, the correction object is the probe angle; if the first round of candidate command actions is an amplification gain action, the correction object is the gain ratio. The first round of candidate command actions are replaced by the first round of actual equivalent actions as the action input for the state evolution process. The target Dreamer world model starts from the initial settlement state and predicts the first predicted settlement state after the first round of actual equivalent actions. When predicting the first predicted settlement state, the target Dreamer world model does not use the first round of candidate command actions as action input, but uses the first round of actual equivalent actions as action input. The first predicted settlement state is the state representation generated by the target Dreamer world model in the latent space, which is used to characterize the leaky light observation stable state expected to be formed after the execution of the first round of actual equivalent actions. Using the first predicted settlement state as the starting point for the second round of deduction, the target Dreamer world model generates the second round of candidate command actions. The execution amount is then corrected again based on the action execution deviation to obtain the second round of actual equivalent actions. These second round of actual equivalent actions are used as the action input for the state evolution process to predict the second predicted settlement state. The second round of deduction follows the processing method of the first round, with the second round of candidate command actions undergoing execution amount correction before entering the state evolution process. The first predicted settlement state is not used as actual sampled data but as the starting point for deduction within the potential space. The second predicted settlement state is not directly output as a detection result but is used to further form candidate adjustment paths. Candidate command actions, actual equivalent actions, and predicted settlement states are continuously generated in the manner described above, forming a candidate adjustment path consisting of several predicted settlement states connected sequentially. The derivation length of the candidate adjustment path can be set to 5 steps. Each candidate adjustment path consists of the initial settlement state, each round of actual equivalent actions, and each round of predicted settlement states in the derivation order. If a candidate adjustment path reaches a preset matching criterion during the derivation process, the expansion of the candidate adjustment path is stopped. In each round of latent space deduction, candidate command actions do not directly participate in the state evolution process. The state evolution process is predicted using the actual equivalent actions of the corresponding round. This process is used to keep the latent space deduction consistent with the action execution deviation written in the training phase. Candidate command actions are only used as the original actions generated by the policy network, and the actual equivalent actions are used as the action inputs for the state evolution process. In this way, the target Dreamer world model predicts the settlement state according to the action magnitude after the actual response during deduction. During the application phase, the model parameters of the target Dreamer world model remain unchanged, and the deviation perception process is only used to correct the action inputs used in the latent space extrapolation.

[0026] In this embodiment, determining the target adjustment action from the remaining candidate adjustment paths specifically includes: Based on the actual equivalent actions, several candidate adjustment paths are generated in the latent space of the target Dreamer world model. Each candidate adjustment path consists of the predicted settlement state and the actual equivalent action connected in the deduction order. The predicted settlement status is read sequentially along each candidate adjustment path. The predicted settlement status is then matched with preset failure criteria to generate a failure identification result for each candidate adjustment path. The preset failure criteria can be set according to the failure conditions of the detection status. The failure conditions of the detection status include: the leakage response corresponding to the predicted settlement status is continuously in the invalid detection range; the signal output corresponding to the predicted settlement status enters the saturation range; the direction determination corresponding to the predicted settlement status is flipped in the continuous deduction steps; the frequency identification result corresponding to the predicted settlement status disappears; or the leakage response corresponding to the predicted settlement status is enhanced but the stability decreases. The failure identification result can be set as a pass flag and a failure flag. When any predicted settlement status meets the failure conditions of the detection status, a failure flag is generated for the corresponding candidate adjustment path. When any predicted settlement state in a candidate adjustment path meets a preset failure criterion, the corresponding candidate adjustment path is truncated from the predicted settlement state that meets the preset failure criterion, and subsequent predicted settlement states are no longer generated. The truncation process can be performed in real time during the derivation of the candidate adjustment path. When the r-th predicted settlement state meets the preset failure criterion, the path segment from the initial settlement state to the r-th predicted settlement state is retained, and the generation process of the (r+1)-th predicted settlement state and subsequent predicted settlement states is stopped. In this way, the target Dreamer world model no longer continues to perform latent space derivation for candidate adjustment paths that have entered a failed state. Candidate adjustment paths that are to be truncated are excluded from the selection range of target adjustment actions. The initial settlement state, actual equivalent action sequence, and failure identification result corresponding to the truncated candidate adjustment path are combined into a pruning template. The pruning template is used to record the state and action combination that caused the candidate adjustment path to fail. The initial settlement state can be represented by the latent state generated by the target Dreamer world model. The actual equivalent action sequence is saved according to the order in which the candidate adjustment paths have been deduced. The failure identification result is used to indicate the failure type that triggered the truncation process. The truncated candidate adjustment path does not participate in the selection of target adjustment actions, but is only used as the basis for path exclusion in the subsequent latent space deduction. When the candidate adjustment path generated by the subsequent latent space deduction has the same initial settlement state range and the same actual equivalent action sequence as the pruning template, the deduction of the subsequent predicted settlement state is stopped; the initial settlement state range can be determined by the latent state distance; when the latent state distance between the initial settlement state of the subsequent candidate adjustment path and the initial settlement state in the pruning template is less than a preset distance threshold, and the actual equivalent action sequence of the first few steps of the subsequent candidate adjustment path is consistent with the actual equivalent action sequence in the pruning template, it is determined that the subsequent candidate adjustment path hits the pruning template; after hitting the pruning template, the target Dreamer world model will no longer continue to generate the subsequent predicted settlement state of the corresponding candidate adjustment path; Among the candidate adjustment paths that have not been truncated, the candidate adjustment path that first meets the preset matching criteria is selected, and the candidate command action corresponding to the first actual equivalent action of the selected candidate adjustment path is determined as the target adjustment action. The preset matching criteria correspond to the matching result label, including the predicted settlement state being within the effective detection range, the predicted settlement state meeting the stable observation requirements, and the direction judgment corresponding to the predicted settlement state being consistent. If multiple candidate adjustment paths that have not been truncated all meet the preset matching criteria, the candidate adjustment path with the fewest deduction steps is selected. If the number of deduction steps is the same, the candidate adjustment path with the smaller fluctuation amplitude of the predicted settlement state is selected. The candidate command action corresponding to the first actual equivalent action of the selected candidate adjustment path is taken as the target adjustment action to be actually executed.

[0027] In this embodiment, outputting the fiber optic interface matching result when the subsequent settlement state meets the preset matching criteria specifically includes: After executing the target adjustment action, leakage observations are re-acquired, and the completion time of the target adjustment action is used as the starting point for subsequent lag settlement. The target adjustment action is determined by the target Dreamer world model after candidate adjustment path deduction and failure trajectory pruning. The target adjustment action acts on the detection parameters of the detection device. After the detection device executes the target adjustment action, leakage observations are re-read from the output end of the PD photodiode or TO photodetector. The completion time is the time when the detection parameter corresponding to the target adjustment action reaches the set value. Subsequent lag settlement starts from the completion time to avoid using transient sampling during the execution of the target adjustment action as the basis for detection judgment. According to the delayed settlement method, the re-acquired light leakage observations are identified as transitional observations and determined as stable observations. The sampling content that is still affected by the target adjustment action and circuit switching process after the settlement starting point is marked as transitional observations. The delayed settlement method is consistent with the delayed settlement method of the initial light leakage observations. The re-acquired light leakage observations are continuously read from the settlement starting point. The short-term fluctuation sampling that occurs immediately after the target adjustment action is completed, the sudden change sampling that occurs after the circuit switching, and the sampling content that has not entered a stable change state are marked as transitional observations. After the transitional observation, when the continuously read leaky light observations enter a stable change state, the corresponding leaky light observations are determined as subsequent stable observations; Filter out transitional observations and encode subsequent stable observations in the state input format of the target Dreamer world model to generate subsequent settlement states; The subsequent settlement status is compared with the preset matching criteria. If the subsequent settlement status meets the preset matching criteria, the fiber optic interface matching result is output. If the subsequent settlement status does not meet the preset matching criteria, the subsequent settlement status is used as the input status for the next round. The preset matching criteria are used to determine whether the leakage light observation obtained by the detection device after target adjustment optimization has reached the judgment state. The preset matching criteria may include the leakage light intensity being within the effective detection range of the detection device, the signal-to-noise ratio meeting the detection requirements, the optical direction determination result being stable, the fluctuation of the optical power estimation value being within the set range, and the amplification circuit not experiencing continuous saturation.

[0028] This invention is applied to non-intrusive service light detection scenarios in fiber optic cable pairing systems. After the fiber under test is inserted into the cable slot of the fiber optic cable pairing system, the clamping mechanism performs an automatic clamping action, causing the fiber under test to form a macrobend in the detection area. Due to macrobend loss at the macrobend position, some of the light transmitted inside the fiber under test leaks outward along the bend. A PD photodiode or TO photodetector located on one side of the bend position receives the leaked light and converts it into an electrical signal, thereby obtaining the leakage light observation. By amplifying, sampling, and calculating the leakage light observation, it is possible to determine whether there is service light in the fiber under test, the transmission direction of the service light, and the magnitude of the service light power.

[0029] In the aforementioned non-intrusive optical detection process, the optical fiber under test does not need to be cut or pulled out. The detection action is completed through the cable tray guide and clamping mechanism. The detection results can be displayed by indicator lights showing the light direction, and the light direction and light power can be displayed synchronously through the APP. For weak light remediation scenarios, after the detection device completes the initial clamping, if the leakage light observation does not meet the preset matching criteria, the target Dreamer world model generates candidate command actions based on the initial settlement state. It then determines the target adjustment action by executing deviation perception, latent space inference, and failure trajectory pruning to adjust the clamping parameters, probe receiving parameters, amplification gain parameters, or frequency identification parameters, so that the subsequently acquired leakage light observations form a stable subsequent settlement state.

[0030] When the subsequent settlement status meets the preset matching criteria, the detection device outputs the fiber optic interface matching result. The fiber optic interface matching result may include the presence status of the service light, the direction of light transmission, and the magnitude of the light power. The target Dreamer world model does not directly output the fiber optic interface matching result, but is used to optimize the detection parameters of the detection device, enabling the detection device to obtain stable and determinable leakage light observations. For weak light remediation scenarios, the above detection results can be used to determine whether weak light exists in the fiber under test, whether the light direction is normal, and whether the light power is within the preset range.

[0031] Example 1: To verify the feasibility of this invention in practice, it was applied to a fiber optic maintenance and testing scenario. In this scenario, the object under test is a pigtail that has been connected to a service. The purpose of the test is to determine whether the fiber optic interface matches the target line without cutting or removing the fiber, and to simultaneously obtain the optical direction and optical power detection results. Existing detection methods typically use fixed clamping force, fixed probe angle, and fixed amplification gain for judgment. When the fiber outer diameter, cladding state, wiring bend state, and service light intensity change, the leakage light signal collected by the detection device is prone to fluctuation. In low-light scenarios, it may not be able to detect stably, while in high-light scenarios, the amplification output may saturate, resulting in unstable optical direction judgment, abrupt changes in optical power readings, and misjudgment of interface matching results.

[0032] In this embodiment, the testing personnel insert the optical fiber under test into the slot of the testing device. The testing device performs non-invasive clamping, causing the optical fiber to form a macrobend in the testing area. Because light leakage occurs at the macrobend, the PD photodiode receives the leaked light and converts it into leakage observation. Unlike traditional methods, this invention does not directly use the immediate sampled values ​​after the clamping action for judgment. Instead, it performs delayed settlement on the initial leakage observation. Immediately after the clamping action, the leakage observation typically contains clamping recovery fluctuations, circuit switching fluctuations, and short-term abrupt sampling. If used directly for judgment, transient changes can easily be mistaken for actual optical power changes. This invention identifies the above-mentioned sampling content as transient observations and, after subsequent continuous sampling enters a stable state, uses the stable portion as the initial settlement state. The data entering the target Dreamer world model has eliminated significant transient interference, more closely resembling the true stable testing state of the testing device under the current optical fiber interface.

[0033] After receiving the initial settlement state, the target Dreamer world model generates candidate command actions for adjusting the detection device. These candidate command actions do not directly participate in state derivation; instead, they are converted into actual equivalent actions through an execution deviation perception process. In actual use, the command actions issued by the detection device do not perfectly match the actual detection results. For example, after a command to fine-tune the tablet, the actual macrobending change may be smaller than the command value due to mechanical hysteresis or elastic compression of the fiber cladding; after a command to change the probe angle, the actual receiving direction change may also deviate from the command value due to assembly deviations; and after a command to switch the amplification gain, the actual output gain will also be affected by the circuit response. This invention incorporates these execution differences into the state evolution process, so that the latent space derivation is no longer based on ideal command actions but on actual equivalent actions, thereby reducing the model's misjudgment of subsequent detection states.

[0034] During the latent space extrapolation process, the target Dreamer world model generates multiple candidate adjustment paths. Each candidate adjustment path corresponds to a possible detection parameter adjustment process. This invention further employs a failure trajectory pruning process to truncate candidate adjustment paths that have shown signs of being unsustainable. The term "unsustainable" does not simply mean a poor result, but rather that continuing to adjust along that path makes it difficult to obtain stable detection results; for example, output saturation trends, repeated changes in light direction judgment, increased fluctuations in light power readings, or persistently unclear weak light characteristics. The truncated paths no longer enter the target adjustment action selection range, reducing ineffective adjustments. The detection device determines the target adjustment action from the remaining candidate adjustment paths and executes that target adjustment action.

[0035] After performing the target adjustment action, the detection device re-acquires leakage light observations and generates subsequent settlement states again using lag settlement. Only when the subsequent settlement states meet the preset matching criteria does the detection device output the fiber interface matching result. This result is not a classification result directly given by the target Dreamer world model, but rather a detection conclusion obtained by the detection device based on stable leakage light observations after parameter optimization. The role of the target Dreamer world model is to adjust the detection parameters of the detection device, improving the quality of sampled data and the stability of the judgment; the fiber interface matching result is output by the detection device based on the optimized stable detection data. This result can be used to determine whether the fiber under test matches the target interface, providing detection results for optical direction and optical power.

[0036] Table 1: Comparison of Non-Invasive Fiber Optic Interface Testing Results

[0037] As can be seen from Table 1, the method of the present invention has significant comprehensive advantages in non-invasive fiber optic interface detection scenarios. Traditional fixed-parameter detection methods are prone to unstable sampling values ​​when the fiber outer diameter, cladding state, and service light intensity change, resulting in insufficient interface matching accuracy and optical direction judgment accuracy.

[0038] This invention eliminates oversampling caused by clamping actions and circuit switching through delayed settlement, making the data entering the judgment process more stable. Therefore, the optical power repeatability deviation is reduced from 0.46dB to 0.13dB. Execution deviation sensing allows the model to use actual equivalent actions during deduction, reducing parameter adjustment errors caused by mechanism response deviations, and improving interface matching accuracy to 97.8%. Failure trajectory pruning reduces invalid adjustment paths, enabling the detection device to determine the target adjustment action more quickly, shortening the average detection time per fiber from 4.0s to 2.6s. For weak light detection, this invention improves stable detection capability under weak light conditions by adaptively adjusting detection parameters, increasing the weak light detection rate from 74.6% to 93.2%.

[0039] Overall, this invention can improve detection stability, reduce false positive rate, and improve fiber optic interface matching efficiency without affecting optical transmission.

[0040] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A fiber optic interface adaptive matching method based on reinforcement learning, characterized in that, Includes the following steps: Collect fiber macrobend leakage light interaction data, annotate the fiber macrobend leakage light interaction data, and generate a training sample set characterizing the fiber interface detection environment. A Dreamer world model is constructed based on the training sample set. Action execution bias, post-action observation lag and failure adjustment trajectory are written into the state evolution process to train the target Dreamer world model. The fiber under test is placed into the detection slot and an initial clamping action is applied. The fiber under test forms a macro-bend leakage state in the detection area, and the initial leakage observation is collected. The initial light leakage observations are settled with lag. Transient observations caused by clamping actions and circuit switching are identified. Transient observations are filtered out to generate stable observations. The initial settlement state is formed by the stable observations. The initial settlement state is input into the target Dreamer world model to generate candidate command actions. The actual equivalent actions corresponding to the candidate command actions are converted by executing the deviation perception process. The latent space inference is based on the actual equivalent actions. Based on the actual equivalent action, candidate adjustment paths are deduced. Candidate adjustment paths that meet the preset failure criteria are identified through the failure trajectory pruning process. The identified candidate adjustment paths are truncated, and the target adjustment action is determined from the remaining candidate adjustment paths. After performing the target adjustment action, the light leakage observation is re-acquired, and the subsequent settlement status is generated according to the lag settlement. When the subsequent settlement status meets the preset matching criteria, the fiber optic interface matching result is output.

2. The fiber optic interface adaptive matching method based on reinforcement learning according to claim 1, characterized in that, The annotation process for the optical fiber macrobend leakage data specifically includes: The training fiber is made to form a macro-bend and light leakage state in the detection slot according to the preset clamping action, and the detection action data and light leakage response data are recorded. Baseline correction, outlier sampling removal, and time alignment are performed on the light leakage response data to obtain the effective light leakage sequence corresponding to the detection action data. Based on the response relationship between the detected action data and the effective light leakage sequence, an action execution deviation label is generated; Based on the magnitude and duration of the changes in the effective light leakage sequence after the detection action, observation settlement labels are generated; Based on whether the detection results corresponding to the valid light leakage sequence meet the preset matching criteria, a matching result label is generated; The detection action data, effective light leakage sequence, action execution deviation annotation, observation settlement annotation, and matching result annotation are combined into sample records in chronological order, and a training sample set representing the fiber optic interface detection environment is generated from the sample records.

3. The fiber optic interface adaptive matching method based on reinforcement learning according to claim 2, characterized in that, The specific steps of incorporating action execution deviation, post-action observation lag, and failure adjustment trajectory into the state evolution process include: Read sample records from the training sample set and encode the valid light leakage sequences in the sample records as the current potential state; The detected action data is combined with the action execution deviation annotation to generate the actual equivalent action, which is then used as the action input for the state evolution process. Based on the observation settlement label, stable observations are extracted from the effective light leakage sequence, and stable observations are encoded as the next potential state. Transitional observations are excluded from the next state in the state evolution process. Based on the matching result annotation and preset failure criteria, a failure adjustment trajectory identifier is generated, and the failure adjustment trajectory identifier is written into the trajectory termination condition of the state evolution process; when generating the failure adjustment trajectory identifier, the matching result annotation in the sample record is read first. The Dreamer world model is trained using the current potential state, the actual equivalent action, the next potential state, and the trajectory termination condition to obtain the target Dreamer world model.

4. The fiber optic interface adaptive matching method based on reinforcement learning according to claim 3, characterized in that, The initial light leakage observation specifically includes: Obtain the clamping start signal after the optical fiber under test is placed into the detection slot, and call the initial action command corresponding to the initial clamping action; An initial sampling task is generated based on the initial action command, and the clamping start time, clamping completion time, and sampling start time are recorded in the initial sampling task; Read the raw light leakage response data of the optical fiber under test after the initial clamping action from the output end of the PD photodiode or TO photodetector. The original light leakage response data is truncated according to the sampling start time to obtain the initial sampling segment; Background baseline subtraction and sample value normalization are performed on the initial sampling segment to generate initial light leakage observations; The initial light leakage observation is linked to the initial action command, the clamping start time, and the clamping completion time to form an initial observation record.

5. The fiber optic interface adaptive matching method based on reinforcement learning according to claim 4, characterized in that, The process of performing delayed settlement on the initial light leakage observation specifically includes: Read the initial light leakage observation, initial action command, clamping start time, and clamping completion time from the initial observation record; The settlement starting point for the initial light leakage observation is taken as the moment the clamping is completed. The sampling content generated by the initial clamping action before the settlement starting point is marked as transitional observation. After the settlement start point, the initial light leakage observations are continuously read in chronological order. Combined with the circuit switching process corresponding to the initial action command, the short-term sudden sampling content caused by the circuit switching is identified and incorporated into the transition observation. The continuity of the remaining initial light leakage observations is judged. When the continuously read sampling content maintains a stable change state relative to the previous sampling content, the corresponding sampling content is determined as a stable observation. If the continuously read sampled content still shows clamping recovery fluctuations or circuit switching fluctuations, it will continue to be marked as a transitional observation and the settlement will be delayed. Transient observations are filtered out, and only stable observations are retained as the state input source for the target Dreamer world model. The initial settlement state is generated based on the stable observations.

6. The fiber optic interface adaptive matching method based on reinforcement learning according to claim 5, characterized in that, The potential space deduction is based on actual equivalent actions and specifically includes: The initial settlement state is input into the target Dreamer world model, which then generates the first round of candidate command actions based on the initial settlement state. Based on the action execution deviation in the state evolution process, the execution amount of the first round of candidate command actions is corrected to obtain the first round of actual equivalent actions. The first round of candidate command actions are replaced by the first round of actual equivalent actions as the action input for the state evolution process. The target Dreamer world model starts from the initial settlement state and predicts the first predicted settlement state after the first round of actual equivalent actions. Using the first predicted settlement state as the starting point for the second round of deduction, the target Dreamer world model generates the second round of candidate command actions. The execution amount is then corrected again based on the action execution deviation to obtain the second round of actual equivalent actions. The second round of actual equivalent actions are used as the action input for the state evolution process to predict the second predicted settlement state. Candidate command actions, actual equivalent actions, and predicted settlement states are continuously generated in the above manner to form a candidate adjustment path that is sequentially connected by several predicted settlement states; In each round of potential space simulation, candidate command actions do not directly participate in the state evolution process; the state evolution process is predicted using the actual equivalent actions of the corresponding round. During the application phase, the model parameters of the target Dreamer world model remain unchanged, and the deviation perception process is only used to correct the action inputs used in the latent space extrapolation.

7. The fiber optic interface adaptive matching method based on reinforcement learning according to claim 6, characterized in that, The specific steps of determining the target adjustment action from the remaining candidate adjustment paths include: Based on the actual equivalent actions, several candidate adjustment paths are generated in the latent space of the target Dreamer world model. Each candidate adjustment path consists of the predicted settlement state and the actual equivalent action connected in the deduction order. The predicted settlement status is read in the order of deduction for each candidate adjustment path, and the predicted settlement status is matched with the preset failure criteria to generate the failure identification result for each candidate adjustment path. When any predicted settlement state in the candidate adjustment path meets the preset failure criterion, the corresponding candidate adjustment path will be truncated from the predicted settlement state that meets the preset failure criterion, and subsequent predicted settlement states after the truncation will no longer be generated. Candidate adjustment paths that are to be truncated are excluded from the selection range of target adjustment actions, and the starting settlement state, actual equivalent action sequence and failure identification result corresponding to the truncated candidate adjustment path are combined into a pruning template; When the candidate adjustment path generated by subsequent potential space deduction has the same initial settlement state range and the same actual equivalent action sequence as the pruning template, stop deducing the subsequent predicted settlement state. The candidate adjustment path that first satisfies the preset matching criterion is selected from the candidate adjustment paths that have never been truncated. The candidate command action corresponding to the first actual equivalent action of the selected candidate adjustment path is determined as the target adjustment action.

8. The fiber optic interface adaptive matching method based on reinforcement learning according to claim 7, characterized in that, The step of outputting the fiber optic interface matching result when the preset matching criterion is met in the subsequent settlement state specifically includes: After the target adjustment action is performed, leakage observations are collected again, and the completion time of the target adjustment action is used as the settlement starting point for subsequent lag settlements. According to the delayed settlement method, the leaky light observations that are re-acquired are identified as transitional observations and stable observations are determined. The sampling content that is still affected by the target adjustment action and circuit switching process after the settlement starting point is marked as transitional observations. After the transitional observation, when the continuously read leaky light observations enter a stable change state, the corresponding leaky light observations are determined as subsequent stable observations; Filter out transitional observations and encode subsequent stable observations in the state input format of the target Dreamer world model to generate subsequent settlement states; The subsequent settlement status is compared with the preset matching criteria. If the subsequent settlement status meets the preset matching criteria, the fiber optic interface matching result is output. If the subsequent settlement status does not meet the preset matching criteria, the subsequent settlement status is used as the input status for the next round.