Optical cable threading adaptive control method and system based on reinforcement learning

Through the reinforcement learning-based adaptive control method for optical cable threading, a multi-source sensor network is used to obtain environmental information and update the model in real time, which solves the problems of low control accuracy and poor stability of optical cable threading in complex environments, and realizes efficient and accurate execution of threading tasks.

CN120595441AInactive Publication Date: 2025-09-05SHENZHEN HUATI AUTOMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511108865.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-09-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing optical cable threading control methods are unable to dynamically respond to environmental changes, resulting in low control accuracy, large threading deviations, unstable equipment operation, and safety hazards in complex environments.

Method used

An adaptive control method for optical cable threading based on reinforcement learning is adopted. Environmental information is obtained through a multi-source sensor network, feature extraction and data fusion are performed, and environmental state characteristics are constructed. A preliminary threading strategy is generated using a pre-trained environmental perception model, and the model is updated in real time during execution to optimize the control strategy.

Benefits of technology

It improves the environmental perception and control adaptability of the optical cable threading system in complex environments, realizes real-time response and precise control of dynamically changing environments, reduces threading deviation and equipment loss, and improves task completion rate and equipment operation stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120595441A_ABST
    Figure CN120595441A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent control, and discloses an optical cable threading adaptive control method and system based on reinforcement learning, and the method comprises the steps: obtaining the first environment information of a to-be-threaded channel, carrying out the feature extraction based on the first environment information, and obtaining the first environment state feature of the to-be-threaded channel; inputting the first environment state feature into a pre-trained environment perception model to obtain a preliminary threading strategy; controlling an optical cable threading device to execute a threading action according to the preliminary threading strategy, and collecting feedback data in the process of executing the threading action; and updating the environment sensing model based on the feedback data, obtaining a second environment state feature of the to-be-threaded channel, and obtaining a target threading strategy based on the second environment state feature and the updated environment sensing model to adjust the threading action of the optical cable threading equipment. According to the method, adaptive control of the optical cable threading equipment can be realized, and the stability and reliability of the threading process are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent control technology, and in particular to a reinforcement learning-based optical cable threading adaptive control method and system. Background Art

[0002] At present, optical cable threading is a key link in the construction of modern communication infrastructure. Its automation and intelligence level directly affects the stability, efficiency, engineering quality and cost control of network transmission. With the rapid expansion of communication networks, achieving efficient and accurate optical cable threading in complex environments has become a key area that the industry urgently needs to break through.

[0003] In one existing technology, optical cable threading control methods often rely on preset rules or manual intervention to plan the threading path and control equipment operations. This existing technology struggles with dynamically responding to the uncertainties introduced by environmental changes during the threading process. This is particularly true in scenarios with varying spatial constraints and complex path geometry. This lack of environmental adaptability and control intelligence limits control accuracy and efficiency, which can easily lead to threading deviations, increased equipment losses, and safety hazards, impacting both equipment operational stability and task completion rates. The core of this problem lies in the inability to achieve adaptive control of the threading equipment based on dynamically acquired environmental information. Summary of the Invention

[0004] The present invention provides a reinforcement learning-based adaptive control method and system for optical cable threading, so as to solve the problem in the prior art that adaptive control of threading equipment cannot be achieved based on dynamically acquired environmental information.

[0005] In a first aspect, in order to solve the above technical problems, the present invention provides an optical cable threading adaptive control method based on reinforcement learning, comprising: Acquire first environmental information of the channel to be threaded, perform feature extraction based on the first environmental information, and obtain first environmental state features of the channel to be threaded; Inputting the first environmental state feature into a pre-trained environmental perception model to obtain a preliminary threading strategy; Controlling the optical cable threading device to perform a threading action according to the preliminary threading strategy, and collecting feedback data during the threading action; The environmental perception model is updated based on the feedback data, and the second environmental state characteristics of the channel to be threaded are obtained. Based on the second environmental state characteristics and the updated environmental perception model, a target threading strategy is obtained to adjust the threading action of the optical cable threading device.

[0006] Preferably, the acquiring of first environmental information of the channel to be threaded and performing feature extraction processing based on the first environmental information to obtain first environmental state features of the channel to be threaded include: Collecting three-dimensional spatial data, obstacle distribution information, and ambient temperature distribution information in the channel to be threaded as the first environmental information based on a pre-deployed multi-source sensor network; performing data fusion processing on the three-dimensional spatial data, the obstacle distribution information, and the ambient temperature distribution information to obtain a first environmental data set; Determining a first geometric feature of each threading path in the to-be-threaded channel based on the first environmental data set; According to each of the first geometric features, a critical path is screened out from all the threading paths and path constraint data corresponding to the critical path is generated; Based on all the path constraint data, the first environmental state feature is integrated and formed.

[0007] Preferably, the step of inputting the first environmental state feature into a pre-trained environmental perception model to obtain a preliminary threading strategy includes: Extracting path change trend features and spatial constraint features based on the first environmental state features; Constructing a state vector according to the path change trend characteristics and the spatial constraint characteristics, and inputting the state vector into the environment perception model to obtain a path action sequence; The preliminary threading strategy is generated according to the path action sequence.

[0008] Preferably, generating the preliminary threading strategy according to the path action sequence includes: determining a path deviation result in a preliminary threading path according to the path action sequence, and judging whether the path deviation result is greater than a preset threshold; If the path deviation result is greater than a preset threshold, obtaining the spatial structural features of the path area corresponding to the path deviation result; Determine the angle variation range of the camera in the channel to be threaded based on the spatial structural characteristics; Setting a viewing angle parameter based on the angle variation range to adjust the camera, and collecting obstacle status information through the adjusted camera; The first environmental state feature is updated based on the obstacle state information, and the step of extracting the path change trend feature and the spatial constraint feature based on the first environmental state feature is returned to, until the path deviation result is less than a preset threshold, to obtain the preliminary threading strategy.

[0009] Preferably, controlling the optical cable threading device to perform the threading action according to the preliminary threading strategy and collecting feedback data during the threading action includes: Generate a control instruction according to the preliminary threading strategy to drive the optical cable threading device to perform a threading action in the to-be-threaded channel according to the preliminary threading path corresponding to the preliminary threading strategy; During the threading operation, the environmental status information around the optical cable threading device is collected based on the multi-source sensor network as the feedback data; The feedback data includes equipment operation trajectory data, obstacle distribution change data and spatial structure change data.

[0010] Preferably, updating the environment perception model based on the feedback data includes: Analyzing the feedback data and constructing an environmental change data set based on the analysis results; Comparing the environmental change data set with a preset environmental parameter benchmark to obtain a difference result; If it is determined that the difference result exceeds a preset range, adjusting the learning parameters in the environment perception model according to the difference result; Through the iterative mechanism of reinforcement learning, the adjusted learning parameters are substituted into the environmental perception model for simulation verification to update the environmental perception model.

[0011] Preferably, the step of obtaining the second environmental state feature of the channel to be threaded includes: During the threading operation, the real-time environmental data around the optical cable threading device is dynamically collected through the adjusted camera; Performing data fusion processing on the real-time environmental data and the feedback data to obtain a second environmental data set, and determining a second geometric feature of the current threading path based on the second environmental data set; The second environmental state feature is obtained according to the second geometric feature.

[0012] Preferably, obtaining a target threading strategy based on the second environmental state feature and the updated environmental perception model to adjust the threading action of the optical cable threading device includes: Inputting the second environmental state feature into the updated environmental perception model to obtain a risk assessment result of the current threading path, and outputting a target action sequence matching the current threading path based on the risk assessment result; The target threading strategy is generated according to the target action sequence, and the target threading strategy is converted into control parameters and sent to the optical cable threading device to adjust the threading action of the optical cable threading device.

[0013] In a second aspect, the present invention provides an optical cable threading adaptive control system based on reinforcement learning, comprising: A data acquisition module is used to acquire first environmental information of the channel to be threaded, and perform multi-source fusion and feature extraction based on the first environmental information to obtain first environmental state features of the channel to be threaded; A strategy processing module is used to input the environmental state characteristics into a pre-trained environmental perception model to obtain a preliminary threading strategy; A feedback collection module, configured to control the optical cable threading device to perform the threading action according to the preliminary threading strategy and to collect feedback data during the threading process; A strategy adjustment module is used to update the environmental perception model based on the feedback data, and obtain the second environmental state characteristics of the channel to be threaded, and based on the second environmental state characteristics and the updated environmental perception model, obtain the target threading strategy to adjust the threading action of the optical cable threading device.

[0014] Preferably, the data acquisition module is further used for: Collecting three-dimensional spatial data, obstacle distribution information, and ambient temperature distribution information in the channel to be threaded as the first environmental information based on a pre-deployed multi-source sensor network; performing data fusion processing on the three-dimensional spatial data, the obstacle distribution information, and the ambient temperature distribution information to obtain a first environmental data set; Determining a first geometric feature of each threading path in the to-be-threaded channel based on the first environmental data set; According to each of the first geometric features, a critical path is screened out from all the threading paths and path constraint data corresponding to the critical path is generated; Based on all the path constraint data, the first environmental state characteristic is determined.

[0015] Compared with the prior art, the present invention has the following beneficial effects: (1) By acquiring the first environmental information of the channel to be threaded and performing feature extraction, it is possible to construct the first environmental state features that conform to the actual conditions of the current threading path, achieve comprehensive perception of multi-dimensional environmental factors such as the path spatial structure and obstacle distribution, and thus improve the adaptability of the threading system to complex environments. The extracted first environmental state features are input into the pre-trained environmental perception model to obtain a preliminary threading strategy, so that the control strategy can be dynamically generated according to the actual environmental state, avoiding the control rigidity caused by relying on static rules, and improving the accuracy and flexibility of control instruction generation; (2) Controlling the optical cable threading equipment to execute the threading action according to the preliminary threading strategy and collecting feedback data during the execution process can achieve closed-loop feedback between the control action and the actual status of the equipment, which helps to timely discover and respond to risk factors caused by environmental changes. By processing the feedback data and updating the environmental perception model, the model parameters can be dynamically optimized, enhancing the model's adaptive learning ability to actual environmental changes, and improving its generalization and robustness; (3) After the model is updated, the second environmental state feature is re-acquired, and the target threading strategy is obtained based on the feature combined with the updated model, so that the control decision can be continuously optimized as the environmental perception capability improves, ensuring that the threading equipment can still maintain high-precision operation under conditions of changing environmental constraints or complex paths, effectively reducing threading deviation, equipment loss and safety risks, and significantly improving the completion rate of threading tasks and the stability of equipment operation; (4) In summary, the present invention provides an adaptive control method for optical cable threading based on reinforcement learning, which can effectively improve the environmental perception and control adaptive capabilities of the optical cable threading system in complex environments, and achieve real-time response and precise control of dynamically changing environments, thereby solving the problems existing in the prior art such as the inability to dynamically respond to environmental changes, low control accuracy, large threading deviation, unstable equipment operation, etc., and improving the automation, intelligence and stability of threading operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 1 is a flow chart of an adaptive control method for optical cable threading based on reinforcement learning provided by the first embodiment of the present invention; Figure 2 2 is a schematic diagram of the structure of an optical cable threading adaptive control system based on reinforcement learning provided by the second embodiment of the present invention. DETAILED DESCRIPTION

[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0018] Reference Figure 1 The first embodiment of the present invention provides an optical cable threading adaptive control method based on reinforcement learning, comprising the following steps: S11, obtaining first environmental information of the channel to be threaded, performing feature extraction based on the first environmental information, and obtaining first environmental state features of the channel to be threaded; S12, inputting the first environmental state feature into a pre-trained environmental perception model to obtain a preliminary threading strategy; S13, controlling the optical cable threading device to execute the threading action according to the preliminary threading strategy, and collecting feedback data during the threading action; S14, updating the environmental perception model based on the feedback data, and obtaining the second environmental state characteristics of the channel to be threaded, and obtaining the target threading strategy based on the second environmental state characteristics and the updated environmental perception model to adjust the threading action of the optical cable threading device.

[0019] In some embodiments, a reinforcement learning-based adaptive control method for optical cable threading is provided to enable dynamic path planning and high-precision control of optical cable threading equipment in complex environments. By dynamically collecting multi-source environmental information, extracting environmental state characteristics, and updating the perception model through reinforcement learning, this method effectively addresses spatial constraint changes and path deviations during the threading process, improving the control accuracy and task completion stability of the threading equipment.

[0020] As a feasible embodiment, a multi-source sensor network is deployed to monitor the threading channel in real time, acquiring a three-dimensional spatial data stream including spatial structure, obstacle distribution, ambient temperature distribution, and channel morphology, which serves as the first environmental information. This first environmental information covers the spatial boundaries and geometric change areas within the threading path, forming a preliminary set of environmental data, namely the first environmental data set. Based on this first environmental data set, dynamic changes in the environment are identified and analyzed, providing basic data support for subsequent feature extraction. Specifically, based on the first environmental information, a data fusion process is used to unify the three-dimensional spatial information from different sources to eliminate redundancy and bias caused by differences in viewing angle, data density, or accuracy. The fusion generates the first environmental data set. Based on this first environmental data set, feature extraction is performed on the spatial constraint boundaries, obstacle locations, and path geometry within the threading channel to construct a first environmental state feature that represents the current channel environmental state. This first environmental state feature includes the locations of key obstacles, path trend change characteristics, and their relationship with the channel boundaries, and can relatively comprehensively reflect the potential risk trends in the current path environment. Subsequently, the first environmental state feature is input into a pre-trained environmental perception model to generate a preliminary threading strategy. This environmental perception model, built through reinforcement learning, predicts potential deviations along the threading path and assesses path risk accordingly. If the model's assessment indicates a high-risk area, it then identifies the necessary viewing angles based on the three-dimensional structure of the environment.

[0021] In a specific embodiment, when the path deviation exceeds a preset threshold, a camera perspective adjustment mechanism is triggered. Based on risk assessment results, this mechanism predicts the optimal perspective switching timing and determines camera angle change parameters that match the current environmental constraints. Based on this, it generates control instructions for the camera's movement trajectory and angle adjustment, which are used for subsequent device perspective control. Next, the camera's angle and position are adjusted according to these camera control instructions, executing the perspective switching operation. After the adjustment, the camera recaptures images and structures of the cable threading path area, generating an updated environmental perception data stream. This data stream contains more accurate spatial features and obstacle distribution information for the current state of the cable threading path, reflecting the latest dynamic environmental conditions. Combined with the updated environmental perception data stream, the path action sequence in the initial cable threading strategy is corrected. If the data stream indicates that the current environmental state has changed beyond the adaptation range of the original model, the cable threading path must be replanned. During the replanning process, the environmental perception model continuously updates learning parameters based on the new environmental characteristics, constructing a new state vector that matches the current environment. Based on this, a target cable threading strategy is generated. This target cable threading strategy is converted into specific control parameters through a control instruction conversion and distribution mechanism and distributed to the optical cable threading equipment. After receiving the control command, the device continues to perform the threading action along the adjusted path. During the threading process, the operating status of the threading equipment is continuously collected, including feedback data such as the actual operating trajectory, path deviation, changes in obstacles and channel boundaries. Based on this feedback data, the environmental perception model continues to be iteratively optimized, and the model's learning weights and policy outputs are cyclically updated to give the model stronger generalization and adaptability, and to be able to adaptively adjust the operation strategy during subsequent threading processes. If it is detected that the control accuracy still does not meet the task requirements after multiple rounds of iterative feedback, the strategy will continue to be adjusted based on the feedback data until the accuracy requirements are met.

[0022] In summary, the present invention's adaptive control method for optical cable threading, based on a reinforcement learning mechanism, achieves efficient operation and adaptive control of optical cable threading equipment in complex spatial environments through dynamic acquisition and feature extraction of first environmental information, real-time assessment of path risk, automatic adjustment of camera viewing angles, cyclical updates of environmental perception models, and the output of a final control strategy. This implementation effectively addresses the existing issues of poor environmental adaptability, large path deviations, and delayed control response during the threading process, providing reliable assurance for the automated and intelligent execution of threading tasks.

[0023] In step S11, first environmental information of the channel to be threaded is obtained, and feature extraction processing is performed based on the first environmental information to obtain first environmental state features of the channel to be threaded, including: Collecting three-dimensional spatial data, obstacle distribution information, and ambient temperature distribution information in the channel to be threaded as first environmental information based on a pre-deployed multi-source sensor network; Performing data fusion processing on the three-dimensional spatial data, the obstacle distribution information, and the ambient temperature distribution information to obtain a first environmental data set; Determining a first geometric feature of each threading path in the threading channel based on the first environmental data set; According to each first geometric feature, a critical path is screened out from all threading paths and path constraint data corresponding to the critical path is generated; Based on all path constraint data, a first environmental state feature is formed by integration.

[0024] In some embodiments, a method for acquiring a first environmental state characteristic of a channel to be threaded is provided. This method deploys a multi-source sensor network to collect multi-dimensional information from the threading environment, fuses the acquired data, and extracts path characteristics. This method identifies key path segments and performs constraint modeling, thereby generating a complete and accurate first environmental state characteristic, providing a data foundation for the subsequent generation of a threading control strategy.

[0025] As a feasible embodiment, first environmental information of the channel to be threaded is obtained through a pre-deployed multi-source sensor network. The multi-source sensor network may include lidar, infrared sensors, ultrasonic sensors, etc., which respectively collect three-dimensional spatial structure, obstacle distribution information, and ambient temperature information, and together constitute the first environmental information. The three-dimensional spatial data primarily reflects the geometric structure of the path to be threaded, including information such as channel width, height, and curvature. Obstacle distribution information is used to identify the location and characteristics of potential interference during the threading process. Ambient temperature information helps identify high-risk areas or abnormal environments that affect equipment stability.

[0026] Based on the first environmental information, a data fusion processing method is used to integrate data from different sources to obtain a first environmental data set with a unified structure. Data fusion processing includes operations such as spatial alignment, coordinate standardization, and redundancy elimination to ensure the integrity and consistency of the data in the first environmental data set and eliminate the deviation caused by the heterogeneity of multiple sensors. For example, after the point cloud data and obstacle detection data are fused, a three-dimensional channel model can be constructed, and the obstacle areas and temperature anomaly areas can be marked. Next, based on the first environmental data set, a spatial structure analysis is performed on each path segment in the threading channel to be processed, and the first geometric features of each path segment are extracted. Geometric features include the width, height, curvature, relative position of obstacles, etc. of the path segment, which are used to characterize the passability and structural stability of the path. For different path segments, evaluation indicators can be set to identify path segments with excessive spatial constraints or drastic geometric changes.

[0027] As a specific embodiment, during the path feature extraction process, the channel to be threaded can be divided into multiple continuous path segments, and the key geometric parameters of each path segment can be calculated. If the width, curvature or obstacle interference level of a certain path segment exceeds a preset threshold, the path segment is marked as a critical path, and its corresponding path constraint data is extracted. The path constraint data may include the minimum passable width of the path segment, the obstacle boundary range, the structural change trend, etc. In addition, based on the path constraint data of all critical path segments, a first environmental state feature can be further formed. The environmental state feature reflects the complexity and risk level of the current threading environment by integrating multi-dimensional parameters such as spatial structure changes, obstacle distribution and temperature information. For example, in an industrial environment, if a certain path segment is detected as a temporary obstacle area, a dynamic change label can be added to the state feature for timely response by subsequent control strategies.

[0028] For example, by deploying lidar to obtain point cloud data, ultrasonic sensors to monitor obstacle distances, and infrared sensors to obtain thermal distribution maps, a first environmental data set containing the complete three-dimensional structure and real-time status of the channel is ultimately formed. After geometric analysis of the data set, it was identified that the width of the third path segment was smaller than the standard, there was a bending area, and it was close to the heat source of the equipment operation, and was determined to be a critical path. The geometric parameters and thermal distribution boundaries of the path segment were extracted to form the path constraint data for this segment. For example, when optimizing a path segment, if the sharp bends or obstacle-dense areas of the original path affect the passability of the equipment, the path reconstruction strategy can be used to recalculate the threading path and update the path constraint data for this segment. The geometric differences before and after optimization are incorporated into the state characteristics to ensure the effectiveness of the subsequent control strategy.

[0029] In summary, through the steps described in this embodiment, including multi-source information collection, data fusion processing, path geometry feature extraction, and critical path screening, it is possible to effectively acquire the first environmental state characteristics covering the entire threading path. This characteristic data not only provides a complete understanding of the current channel environment but also provides a quantifiable and adjustable path information foundation for the threading equipment to execute threading operations in various complex scenarios, facilitating the dynamic generation and adaptive optimization of path control strategies.

[0030] In step S12, the first environmental state feature is input into the pre-trained environmental perception model to obtain a preliminary threading strategy, including: Extracting the path change trend feature and the spatial constraint feature based on the first environmental state feature; A state vector is constructed based on the path change trend characteristics and spatial constraint characteristics, and the state vector is input into the environment perception model to obtain the path action sequence; Generate a preliminary threading strategy based on the path action sequence.

[0031] In some embodiments, the method analyzes the first environmental state characteristics, extracts the path change trend characteristics and spatial constraint characteristics, and based on the constructed state vector, outputs the path action sequence with the help of a pre-trained environmental perception model, thereby generating a preliminary threading strategy.

[0032] As a feasible embodiment, the path change trend feature is used to characterize dynamic fluctuations along the path, such as the shrinking of channel width and the frequent appearance of obstacles. The spatial constraint feature is used to describe the static constraints imposed by the environment on path feasibility, such as path width, clearance height, viewing angle, and minimum safe distance. These two feature types are identified through feature extraction operations to construct a state vector for input.

[0033] In one specific embodiment, the state vector is constructed based on parameters such as the local width change rate of the channel, the degree of curvature change, the frequency of obstacles, the effective area ratio of the traversable area, the obstacle density index at key points along the path, and the minimum required turning angle range for the path. After normalization and encoding, the state vector is input into the environmental perception model. This model is a pre-trained deep neural network structure, trained on data from multiple typical wire-threading scenarios. It learns path adjustment strategies by mimicking the decision-making process of human path planning.

[0034] As a specific embodiment, the environment perception model outputs a path action sequence based on the input state vector. This path action sequence is an ordered set of actions that describe the appropriate travel behavior under the current path state. These actions include maintaining the current direction, fine-tuning the travel angle, adjusting the posture on the spot, and skipping the current path segment. This path action sequence reflects the system's comprehensive assessment of the path state and provides a basis for subsequent path planning and execution control.

[0035] For example, in a certain wire threading scenario, the first environmental state feature shows that the width of the third path channel suddenly drops from 2.0 meters to 1.2 meters, and there is also an obstruction caused by dynamically stacked materials. The path change trend feature extracted is "sharp drop in width", and the spatial constraint feature is "the existence of a blocking area and the minimum turning angle is limited." Based on this, a state vector is constructed and input into the environmental perception model. The path action sequence output by the model includes action instructions such as "perform local adjustment", "detour 15 degrees to the right", and "enter the backup channel". Based on this path action sequence, a preliminary wire threading strategy is constructed to guide the equipment to select a feasible path and avoid potential obstacles.

[0036] It's important to note that by converting path trend characteristics and spatial constraint features into state vectors and inputting them into an environmental perception model to obtain a path action sequence, the system can perceive and respond to complex environments and effectively adapt to the changing demands of threading tasks. This initial threading strategy provides the foundation for subsequent path refinement, camera adjustment, and device command generation, forming the core of the entire adaptive path planning system.

[0037] In step S12, a preliminary threading strategy is generated based on the path action sequence, including: Determine the path deviation result in the preliminary threading path according to the path action sequence, and judge whether the path deviation result is greater than a preset threshold; If the path deviation result is greater than a preset threshold, the spatial structure feature of the path area corresponding to the path deviation result is obtained; Determine the angle variation range of the camera in the channel to be threaded based on the spatial structural characteristics; Setting viewing angle parameters based on the angle variation range to adjust the camera, and collecting obstacle status information through the adjusted camera; The first environmental state feature is updated based on the obstacle state information, and the step of extracting the path change trend feature and the spatial constraint feature based on the first environmental state feature is returned to, until the path deviation result is less than the preset threshold, and a preliminary threading strategy is obtained.

[0038] In some embodiments, the generation of a preliminary threading strategy is based on the path action sequence output by the environmental perception model. Specifically, this includes steps such as determining path deviation, adjusting the camera's viewing angle, collecting obstacle status information, and updating the first environmental state characteristics. This process aims to verify and correct the feasibility of path execution in a dynamic environment, and iteratively update the strategy as necessary, ultimately forming a preliminary threading strategy that satisfies the path deviation constraints.

[0039] As a feasible embodiment, the path deviation result in the preliminary threading path is determined based on the path action sequence. The path deviation result is used to describe the difference between the actual traffic capacity of the path and the planned path, and the gap between the channel width, traffic height, turning angle, etc. and the safety threshold is often used as a measurement indicator. In actual applications, the acquisition of path deviation can be calculated in combination with the real-time sensor data and the path behavior record. When the path deviation result is greater than the preset threshold, it means that the current path does not meet the safe passage conditions and needs to enter the correction process. When the path deviation exceeds the preset threshold, the spatial structure characteristics of the deviation path area are further obtained. The spatial structure characteristics are a data set that describes the static or dynamic constraints in the local area of ​​the path, including but not limited to parameters such as local channel geometry, obstacle density, passable area morphology, turning radius, obstacle distribution and size. These features are collected and processed by multi-source sensing devices to ensure that they reflect the actual physical structure of the current path area.

[0040] As a specific embodiment, the angle variation range of the camera in the channel to be threaded is determined based on the acquired spatial structural features. The angle variation range is expressed as an optional interval of the pitch angle, yaw angle and field of view angle that can be adjusted by the camera, aiming to achieve optimal coverage of the current obstacle area. This step is mainly used to guide the viewing angle adjustment operation and provide conditions for subsequent acquisition of more accurate obstacle status information. For example, in a certain threading path scenario, the path action sequence guides the optical cable threading equipment into a channel area where materials are stacked. The sensor obtains that the channel width is reduced from 1.8 meters to 1.0 meters. The deviation result is determined to exceed the preset threshold of 0.5 meters, and the viewing angle adjustment needs to be triggered. For example, based on the environmental model data, it is known that the obstacle height is 1.2 meters and the width is 0.8 meters. The adjustment angle range can be determined to be +10° to +25° for the pitch angle and -15° to +15° for the yaw angle according to the camera setting position.

[0041] Subsequently, specific viewing angle parameters are set based on the angle variation range to adjust the camera. Viewing angle parameters include, but are not limited to, the center position of the field of view, rotation angle, and zoom level to ensure that the camera's field of view covers the area currently containing obstacles. After adjustment, the camera captures images of the target area to obtain obstacle status information, including its location, size, morphological changes, and motion trends. Based on the obstacle status information captured by the camera, the first environmental state feature is updated. The updated first environmental state feature, combined with the newly collected data, reflects the true state of the current path environment, including the latest spatial layout, obstacle changes, and adjustments to the traversable area. This first environmental state feature is re-input into the path feature extraction and environmental perception model process to extract updated path variation trend features and spatial constraint features, and regenerate the state vector. After the updated state vector is re-input into the environmental perception model, the system obtains a new path action sequence and continues to determine whether the path deviation results meet the preset requirements. If not, the spatial structure feature extraction, camera adjustment, and state update operations are repeated. The process continues to iterate in a circular manner until the path deviation result is less than the preset threshold, at which point the preliminary threading strategy that finally meets the conditions is obtained.

[0042] It should be noted that the method described in this embodiment implements a linkage mechanism between path deviation identification and camera perspective adjustment, dynamically updates perception features and optimizes path strategies based on real-time changes in the environment, effectively improving the adaptability and stability of path execution of the threading equipment in complex environments, and providing a solid foundation for subsequent refined path planning and instruction generation.

[0043] In step S13, the optical cable threading device is controlled to perform the threading action according to the preliminary threading strategy, and feedback data is collected during the threading action, including: Generate control instructions based on the preliminary threading strategy to drive the optical cable threading device to perform threading actions in the channel to be threaded according to the preliminary threading path corresponding to the preliminary threading strategy; During the threading process, the environmental status information around the optical cable threading equipment is collected based on the multi-source sensor network as feedback data; Among them, the feedback data includes equipment operation trajectory data, obstacle distribution change data and spatial structure change data.

[0044] In some embodiments, control instructions are generated based on the preliminary threading strategy to drive the optical cable threading device to perform threading actions in the channel to be threaded according to the preliminary threading path corresponding to the preliminary threading strategy. The preliminary threading strategy is determined based on the path action sequence generated by the environmental perception model, and includes planning information of the threading path and equipment operating parameters, etc. These abstract strategy information are converted into specific operating parameters that can be executed by the device, for example, the position coordinates in the path planning are converted into the direction of travel and speed instructions of the device. These control instructions are transmitted to the execution unit of the optical cable threading device through the communication interface, and the device is driven to perform threading actions in the channel to be threaded according to the preliminary threading path. For example, if the preliminary threading strategy requires the device to slow down in a certain path section and offset to the right a certain distance, corresponding speed control instructions and steering instructions are generated to ensure that the device accurately performs the action. Control instructions include control parameters such as motor speed, steering angle, propulsion force, acceleration and deceleration, stopping, and obstacle avoidance.

[0045] In some embodiments, during the threading process, a multi-source sensor network collects environmental status information surrounding the optical cable threading equipment as feedback data. This multi-source sensor network includes various types of sensors distributed around the equipment and the channel to be threaded. These sensors collect real-time environmental data around the equipment, generating feedback data. This feedback data includes equipment trajectory data, obstacle distribution change data, and spatial structure change data. Equipment trajectory data is acquired through the equipment's own positioning system and motion sensors, reflecting the deviation between the equipment's actual path and the initial threading path. Obstacle distribution change data is collected through sensors such as lidar and cameras, recording the real-time position, size, and movement of obstacles in the environment. Spatial structure change data is obtained through real-time monitoring of environmental geometric features, such as changes in channel width and height, and the need for dynamic adjustment of path curvature. For example, during the threading process, a lidar scans the surrounding environment in real time, detecting new obstacles in a certain area, and a camera captures changes in object outlines caused by changes in light. Together, these data constitute feedback data, providing a basis for subsequent path optimization.

[0046] While the equipment is threading, each sensor continuously collects data at a preset sampling frequency. This data is then preprocessed, including filtering, calibration, and fusion, to improve its accuracy and reliability. This preprocessed data is categorized into equipment trajectory data, obstacle distribution change data, and spatial structure change data, and stored in a feedback database. For example, point cloud data collected by the lidar is filtered to remove noise points and then fused with image data collected by the camera to generate accurate obstacle distribution change data. This feedback data is transmitted in real time to the path optimization module to evaluate the effectiveness of the current threading operation and provide a basis for subsequent path adjustments.

[0047] For example, in a cable threading scenario, the optical cable threading equipment performs threading actions in the channel according to the preliminary threading strategy. When collecting feedback data based on the multi-source sensor network, the lidar detects that the width of the channel ahead has narrowed due to the temporary stacking of materials, and the camera captures the specific position and size of the materials. These data are collected as obstacle distribution change data. At the same time, the positioning system of the optical cable threading equipment records the actual operation trajectory, compares it with the preliminary threading path, and generates equipment operation trajectory data. If it is detected that the channel width change exceeds the preset threshold, the path will be replanned based on this feedback data, and the control instructions will be adjusted to ensure that the equipment passes through the area safely.

[0048] In summary, this embodiment forms a complete threading execution and feedback mechanism by converting a preliminary threading strategy into control instructions to drive the device to execute the threading action. It also utilizes a multi-source sensor network to collect real-time feedback data, including device trajectory data, obstacle distribution change data, and spatial structure change data. This mechanism ensures that the device can execute the threading action according to the planned path in complex environments, while also sensing environmental changes in real time, providing data support for subsequent path optimization and adaptive adjustments, thereby improving the success rate and efficiency of threading tasks.

[0049] In step S14, the environment perception model is updated based on the feedback data, including: Analyze the feedback data and build an environmental change dataset based on the analysis results; Compare the environmental change dataset with the preset environmental parameter benchmark to obtain the difference results; If the difference result is judged to be beyond the preset range, the learning parameters in the environment perception model are adjusted according to the difference result; Through the iterative mechanism of reinforcement learning, the adjusted learning parameters are substituted into the environmental perception model for simulation verification to update the environmental perception model.

[0050] In some embodiments, the feedback data is parsed, and an environmental change data set is constructed based on the parsing results. The parsing process is to structure the multi-dimensional feedback data and extract key information that can reflect environmental changes. For example, the deviation value from the preliminary threading path is extracted from the equipment operation trajectory data, the position, size and movement trend of the new obstacle is parsed from the obstacle distribution change data, and the change in geometric parameters such as channel width and height is obtained from the spatial structure change data. By integrating and classifying these key information, an environmental change data set is constructed, which describes the changes in the current environment relative to the preset environmental parameter benchmark in a quantitative form.

[0051] As a feasible embodiment, the environmental change data set is compared with the preset environmental parameter benchmark to obtain a difference result. The preset environmental parameter benchmark is determined during the initialization phase or the last model update, and represents the various parameter indicators under the ideal environmental state. The comparison process is to compare the various parameters in the environmental change data set with the preset benchmark dimension by dimension, and calculate the difference value between the two. For example, the current channel width is compared with the preset standard width to obtain the difference value of the width change, and the actual position of the obstacle is compared with the preset obstacle-free area to determine the degree of intrusion of the obstacle. These difference values ​​constitute the difference result, which is used to evaluate the degree of deviation between the current environment and the expected environment.

[0052] As a specific embodiment, if the difference result is judged to be beyond the preset range, the learning parameters in the environmental perception model are adjusted according to the difference result. The preset range is a threshold interval pre-set according to the task requirements, which is used to determine whether the environmental changes are significant enough to require updating the model. When one or more parameters in the difference result exceed the preset range, it means that the current environment has changed significantly, and the original environmental perception model may no longer be applicable. At this time, the learning parameters in the environmental perception model are adjusted according to the specific circumstances of the difference result. For example, if it is found that the moving speed of the obstacle exceeds the preset range, the sensitivity parameters of the model to dynamic obstacles may be increased. If the change in channel width makes path planning difficult, the weight parameters related to spatial constraints are adjusted. These adjustments are intended to enable the model to better adapt to new environmental conditions.

[0053] As a specific embodiment, the adjusted learning parameters are substituted into the environmental perception model through the iterative mechanism of reinforcement learning for simulation verification to update the environmental perception model. Reinforcement learning is a machine learning method that learns the optimal strategy by interacting with the environment and the rewards obtained based on feedback. After the adjusted learning parameters are substituted into the environmental perception model, the model is verified using historical feedback data and test scenarios generated by simulation. The model performs threading actions in a simulated environment and obtains corresponding reward values ​​based on the execution results. The reward value is designed based on the task objectives, such as successfully avoiding obstacles, keeping the path deviation within the allowable range, etc. Through multiple iterative training, the model continuously optimizes the learning parameters until the performance in the simulated environment meets the preset performance indicators. At this point, the adjusted model is considered to be a model that is more suitable for the current environment, and the update of the environmental perception model is completed.

[0054] For example, in a wire-threading scenario, an environmental change dataset constructed based on feedback data shows that the number of obstacles in a certain area has increased significantly and their movement speed has accelerated, with the difference from the preset environmental parameter baseline exceeding the preset range. Based on this difference, the learning parameters related to obstacle detection and path planning in the environmental perception model are adjusted. Then, through the iterative mechanism of reinforcement learning, the adjusted model is verified in a simulated environment. Through multiple attempts, the model learns how to better cope with the increase in the number of obstacles and the increase in movement speed, such as planning detours in advance and increasing the time window for obstacle prediction. After verification, the model's performance has been improved, and it can more accurately predict path risks and generate reasonable wire-threading strategies, completing the update of the environmental perception model.

[0055] In summary, this embodiment forms a complete environment-awareness model update mechanism through analysis of feedback data, comparison of differences, parameter adjustment, and reinforcement learning verification. This mechanism dynamically adjusts model parameters based on real-time environmental changes, ensuring that the environment-awareness model remains adaptable to the current environment. This improves the intelligent decision-making capabilities of optical cable threading equipment in complex environments and enhances the efficiency of threading tasks.

[0056] In step S14, the second environmental state characteristics of the channel to be threaded are obtained, including: During the threading process, the real-time environmental data around the optical cable threading equipment is dynamically collected through the adjusted camera; Performing data fusion processing on the real-time environmental data and the feedback data to obtain a second environmental data set, and determining a second geometric feature of the current threading path based on the second environmental data set; A second environmental state feature is obtained according to the second geometric feature.

[0057] In some embodiments, during the execution of the threading action, the real-time environmental data around the optical cable threading device is dynamically collected through an adjusted camera. The parameters of the adjusted camera are set according to the angle change range determined by the spatial structural characteristics of the path area corresponding to the path deviation result. Its viewing angle can cover the current position of the threading device and the surrounding key areas to ensure that the dynamic changes of the environment are fully captured. The real-time environmental data collected by the camera includes but is not limited to the location, shape, movement trajectory of obstacles, the geometric outline of the channel, and light conditions. These data are transmitted to the data processing module in real time in the form of images or video streams, providing a basis for subsequent environmental feature extraction.

[0058] As a feasible embodiment, data fusion processing is performed on the real-time environmental data and feedback data to obtain a second environmental data set, and based on the second environmental data set, the second geometric features of the current threading path are determined. Data fusion processing is the process of integrating and complementing the real-time environmental data collected by the camera with the feedback data. Through technical means such as data alignment, time synchronization and information complementation, redundancy and contradictions between data are eliminated to form a unified, accurate and complete second environmental data set. Based on the second environmental data set, the geometric properties of the current threading path are analyzed, including the length, curvature, width change, branching, etc. of the path, to form a second geometric feature. For example, by analyzing continuous frame images, the curvature of the path and the width change trend are calculated, and the passability of the path is verified in combination with the equipment operation trajectory data.

[0059] As a specific embodiment, a second environmental state feature is obtained based on the second geometric feature. The second environmental state feature is a comprehensive description of the current environment of the channel to be threaded, which is constructed based on the second geometric feature and in combination with the physical properties and constraints of the environment. Through further analysis and quantification of the second geometric feature, key indicators that can reflect the complexity of the environment, the degree of constraint and the trend of change are extracted, such as the tortuosity of the path, the spatial constraint coefficient, the obstacle density, etc. These indicators together constitute the second environmental state feature, which provides an accurate description of the environment for subsequent environmental perception and threading strategy optimization. For example, if the second geometric feature shows that there are multiple sharp turns in the path and the channel width is narrowed, the corresponding tortuosity and spatial constraint coefficient in the second environmental state feature will be higher, indicating that a more cautious threading strategy is needed.

[0060] In summary, this embodiment forms a complete environmental feature acquisition mechanism by dynamically collecting real-time environmental data through an adjusted camera, fusing it with feedback data, extracting secondary geometric features, and constructing secondary environmental state features. This mechanism ensures that the system can accurately perceive environmental changes in the channel to be threaded in real time, providing reliable data support for updating the environmental perception model and optimizing the threading strategy, thereby improving the adaptability of optical cable threading equipment in complex environments and the success rate of threading tasks.

[0061] In step S14, based on the second environmental state characteristics and the updated environmental perception model, a target threading strategy is obtained to adjust the threading action of the optical cable threading device, including: Inputting the second environmental state feature into the updated environmental perception model to obtain a risk assessment result of the current threading path, and outputting a target action sequence matching the current threading path based on the risk assessment result; A target threading strategy is generated according to the target action sequence, and the target threading strategy is converted into control parameters and sent to the optical cable threading device to adjust the threading action of the optical cable threading device.

[0062] In some embodiments, the second environmental state feature is input into the updated environmental perception model to obtain the risk assessment result of the current threading path, and the target action sequence that matches the current threading path is output based on the risk assessment result. After the second environmental state feature is input into the model, the model evaluates the risks of the current threading path, such as collision risk, path blockage risk, etc., by analyzing indicators such as the tortuosity of the path, the spatial constraint coefficient, and the obstacle density, and generates corresponding risk assessment results. Based on the risk assessment results, the model selects the optimal action combination from the predefined action set to form a target action sequence that matches the current threading path. This sequence includes specific actions that the threading device should perform in different positions and environmental conditions, such as forward, backward, turning, adjusting speed, etc.

[0063] As a feasible embodiment, a target threading strategy is generated based on the target action sequence, and the target threading strategy is converted into control parameters and sent to the optical cable threading equipment to adjust the threading action of the optical cable threading equipment. The target threading strategy is a further optimization and integration of the target action sequence. It takes into account the timing relationship between actions, the kinematic constraints of the equipment, and the overall goal of the task. By converting the target action sequence into a specific motion trajectory planning, control parameter setting and time series arrangement, a complete target threading strategy is formed. The target threading strategy is converted into control parameters such as motor speed, steering angle, propulsion force, etc., and sent to the control system of the optical cable threading equipment through a communication interface. The optical cable threading equipment adjusts its own threading action according to the received control parameters, achieving real-time response and adaptation to environmental changes. For example, if the target threading strategy requires the equipment to slow down and perform fine operations in a certain area, the control parameters will adjust the propulsion speed and operation accuracy of the equipment accordingly to ensure the smooth progress of the threading task.

[0064] As a specific example, when generating a target threading strategy, the system also considers the feasibility and safety of the action. By analyzing the device's kinematic model and environmental constraints, the system verifies the executable nature of the target action sequence, avoiding actions that exceed the device's capabilities or could cause collisions. Furthermore, the system evaluates the safety of the strategy to ensure that the device maintains a safe distance from obstacles during execution, mitigating operational risks. For example, in areas with narrow paths and dense obstacles, the system prioritizes conservative action plans to increase safety margins.

[0065] For example, in a scenario involving multiple intersecting pipes, the second environmental state feature indicates the presence of a moving obstacle at an intersection with significant spatial constraints. This feature is fed into the updated environmental perception model, which assesses the high collision risk for this path segment and outputs a target action sequence consisting of waiting, avoiding, and precise steering. The target threading strategy generated based on this sequence instructs the device to pause, wait for the obstacle to pass, and then turn at a precise angle to enter the target pipe. After the control parameters are sent to the device, it accurately executes the strategy, successfully avoiding the obstacle and completing the threading task.

[0066] This embodiment integrates the second environmental state characteristics into an updated environmental perception model to perform risk assessment and action sequence generation. This is then converted into an executable target threading strategy and used to control device execution, forming a complete threading action adjustment mechanism. This mechanism dynamically optimizes threading strategies based on real-time environmental changes, improving the device's adaptability in complex environments and the success rate of threading tasks, ensuring efficient and safe cable threading.

[0067] In summary, the present invention discloses an adaptive control method for optical cable threading based on reinforcement learning, which can effectively improve the environmental perception and control adaptability of the optical cable threading system in complex environments, and realize real-time response and precise control of dynamically changing environments, thereby solving the problems existing in the prior art such as the inability to dynamically respond to environmental changes, low control accuracy, large threading deviation, unstable equipment operation, etc., and improving the automation, intelligence and stability of threading operations.

[0068] Reference Figure 2 The second embodiment of the present invention provides an optical cable threading adaptive control system based on reinforcement learning, comprising: A data acquisition module is used to acquire first environmental information of the channel to be threaded, and perform multi-source fusion and feature extraction based on the first environmental information to obtain first environmental state features of the channel to be threaded; The strategy processing module is used to input the environmental state characteristics into the pre-trained environmental perception model to obtain the preliminary threading strategy; A feedback collection module is used to control the optical cable threading equipment to perform the threading action according to the preliminary threading strategy and collect feedback data during the threading process; The strategy adjustment module is used to update the environmental perception model based on feedback data and obtain the second environmental state characteristics of the channel to be threaded. Based on the second environmental state characteristics and the updated environmental perception model, the target threading strategy is obtained to adjust the threading action of the optical cable threading equipment.

[0069] It should be noted that the optical cable threading adaptive control system based on reinforcement learning provided in an embodiment of the present invention is used to execute all the process steps of the optical cable threading adaptive control method based on reinforcement learning in the above embodiment. The working principles and beneficial effects of the two correspond one to one, so they will not be repeated here.

[0070] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.

[0071] The electronic device may be a computing device such as a desktop computer, notebook, PDA, or smart tablet. The electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will appreciate that the aforementioned components are merely examples of electronic devices and do not constitute a limitation of the electronic device. The electronic device may include more or fewer components than those described above, or a combination of certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, and the like.

[0072] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. The processor is the control center of the electronic device and connects various parts of the entire electronic device using various interfaces and lines.

[0073] The memory can be used to store the computer programs and / or modules. The processor implements the various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory and accessing the data stored in the memory. The memory may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as a sound playback function or an image playback function); the data storage area may store data generated based on the use of the mobile phone (such as audio data, a phone book, etc.). Furthermore, the memory may include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0074] If the module / unit integrated into the electronic device is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention can implement all or part of the process steps in the above-mentioned method embodiments by using a computer program to instruct the relevant hardware. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium can be appropriately increased or decreased based on the requirements of legislation and patent practice in a jurisdiction. For example, in some jurisdictions, based on legislation and patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0075] It should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art can understand and implement the present invention without inventive effort.

[0076] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A reinforcement learning-based adaptive control method for optical cable threading, characterized in that: include: Acquire first environmental information of the channel to be threaded, perform feature extraction based on the first environmental information, and obtain first environmental state features of the channel to be threaded; Inputting the first environmental state feature into a pre-trained environmental perception model to obtain a preliminary threading strategy; Controlling the optical cable threading device to perform a threading action according to the preliminary threading strategy, and collecting feedback data during the threading action; The environmental perception model is updated based on the feedback data, and the second environmental state characteristics of the channel to be threaded are obtained. Based on the second environmental state characteristics and the updated environmental perception model, a target threading strategy is obtained to adjust the threading action of the optical cable threading device.

2. The optical cable threading adaptive control method based on reinforcement learning according to claim 1 is characterized in that: The step of obtaining first environmental information of the channel to be threaded and performing feature extraction based on the first environmental information to obtain first environmental state features of the channel to be threaded includes: Collecting three-dimensional spatial data, obstacle distribution information, and ambient temperature distribution information in the channel to be threaded as the first environmental information based on a pre-deployed multi-source sensor network; performing data fusion processing on the three-dimensional spatial data, the obstacle distribution information, and the ambient temperature distribution information to obtain a first environmental data set; Determining a first geometric feature of each threading path in the to-be-threaded channel based on the first environmental data set; According to each of the first geometric features, a critical path is screened out from all the threading paths and path constraint data corresponding to the critical path is generated; Based on all the path constraint data, the first environmental state feature is integrated and formed.

3. The optical cable threading adaptive control method based on reinforcement learning according to claim 2 is characterized in that: The step of inputting the first environmental state feature into a pre-trained environmental perception model to obtain a preliminary threading strategy includes: Extracting path change trend features and spatial constraint features based on the first environmental state features; Constructing a state vector according to the path change trend characteristics and the spatial constraint characteristics, and inputting the state vector into the environment perception model to obtain a path action sequence; The preliminary threading strategy is generated according to the path action sequence.

4. The optical cable threading adaptive control method based on reinforcement learning according to claim 3 is characterized in that: Generating the preliminary threading strategy according to the path action sequence includes: determining a path deviation result in a preliminary threading path according to the path action sequence, and judging whether the path deviation result is greater than a preset threshold; If the path deviation result is greater than a preset threshold, obtaining the spatial structural features of the path area corresponding to the path deviation result; Determine the angle variation range of the camera in the channel to be threaded based on the spatial structural characteristics; Setting a viewing angle parameter based on the angle variation range to adjust the camera, and collecting obstacle status information through the adjusted camera; The first environmental state feature is updated based on the obstacle state information, and the step of extracting the path change trend feature and the spatial constraint feature based on the first environmental state feature is returned to, until the path deviation result is less than a preset threshold, to obtain the preliminary threading strategy.

5. The optical cable threading adaptive control method based on reinforcement learning according to claim 2, characterized in that: The controlling the optical cable threading device to perform the threading action according to the preliminary threading strategy and collecting feedback data during the threading action includes: Generate a control instruction according to the preliminary threading strategy to drive the optical cable threading device to perform a threading action in the to-be-threaded channel according to the preliminary threading path corresponding to the preliminary threading strategy; During the threading operation, the environmental status information around the optical cable threading device is collected based on the multi-source sensor network as the feedback data; The feedback data includes equipment operation trajectory data, obstacle distribution change data and spatial structure change data.

6. The optical cable threading adaptive control method based on reinforcement learning according to claim 2, characterized in that: The updating of the environment perception model based on the feedback data includes: Analyzing the feedback data and constructing an environmental change data set based on the analysis results; Comparing the environmental change data set with a preset environmental parameter benchmark to obtain a difference result; If it is determined that the difference result exceeds a preset range, adjusting the learning parameters in the environment perception model according to the difference result; Through the iterative mechanism of reinforcement learning, the adjusted learning parameters are substituted into the environmental perception model for simulation verification to update the environmental perception model.

7. The optical cable threading adaptive control method based on reinforcement learning according to claim 4 is characterized in that: The obtaining of the second environmental state feature of the channel to be threaded includes: During the threading operation, the real-time environmental data around the optical cable threading device is dynamically collected through the adjusted camera; Performing data fusion processing on the real-time environmental data and the feedback data to obtain a second environmental data set, and determining a second geometric feature of the current threading path based on the second environmental data set; The second environmental state feature is obtained according to the second geometric feature.

8. The optical cable threading adaptive control method based on reinforcement learning according to claim 7 is characterized in that: The step of obtaining a target threading strategy based on the second environmental state feature and the updated environmental perception model to adjust the threading action of the optical cable threading device includes: Inputting the second environmental state feature into the updated environmental perception model to obtain a risk assessment result of the current threading path, and outputting a target action sequence matching the current threading path based on the risk assessment result; The target threading strategy is generated according to the target action sequence, and the target threading strategy is converted into control parameters and sent to the optical cable threading device to adjust the threading action of the optical cable threading device.

9. An adaptive control system for optical cable threading based on reinforcement learning, characterized in that: include: A data acquisition module is used to acquire first environmental information of the channel to be threaded, and perform multi-source fusion and feature extraction based on the first environmental information to obtain first environmental state features of the channel to be threaded; A strategy processing module is used to input the environmental state characteristics into a pre-trained environmental perception model to obtain a preliminary threading strategy; A feedback collection module, configured to control the optical cable threading device to perform the threading action according to the preliminary threading strategy and to collect feedback data during the threading process; A strategy adjustment module is used to update the environmental perception model based on the feedback data, and obtain the second environmental state characteristics of the channel to be threaded, and based on the second environmental state characteristics and the updated environmental perception model, obtain the target threading strategy to adjust the threading action of the optical cable threading device.

10. The optical cable threading adaptive control system based on reinforcement learning according to claim 9, characterized in that: The data acquisition module is also used for: Collecting three-dimensional spatial data, obstacle distribution information, and ambient temperature distribution information in the channel to be threaded as the first environmental information based on a pre-deployed multi-source sensor network; performing data fusion processing on the three-dimensional spatial data, the obstacle distribution information, and the ambient temperature distribution information to obtain a first environmental data set; Determining a first geometric feature of each threading path in the to-be-threaded channel based on the first environmental data set; According to each of the first geometric features, a critical path is screened out from all the threading paths and path constraint data corresponding to the critical path is generated; Based on all the path constraint data, the first environmental state characteristic is determined.