Control method, device, computer equipment and storage medium of optical fiber locator
By optimizing the rotation path of the optical fiber locator through the target reinforcement learning model, the problems of low optical fiber positioning efficiency and high collision risk in the existing technology are solved, and efficient and safe optical fiber positioning is achieved.
Patent Information
- Application Number
- CN202411017754.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-26
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-07-26
AI Technical Summary
Existing fiber optic positioning methods are inefficient and lack flexibility and adaptability when multiple positioners are densely arranged on the focal plane, and there is a high risk of collision between robotic arms.
The target reinforcement learning model is used to calculate the rotation path of the fiber locator. By determining the robotic arm posture of the fiber locator to be positioned and its adjacent fiber locators, the rotation path is optimized to avoid collision. The training model is used to continuously adjust the strategy to improve positioning efficiency and reduce the probability of collision.
The efficiency of optical fiber positioning is improved, the possibility of collision between robotic arms is reduced, and efficient and safe optical fiber positioning is achieved.
Smart Images

Figure CN118981080B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of optical fiber sensing technology, and in particular to a control method, device, computer equipment, and storage medium for an optical fiber locator. Background Art
[0002] Fiber positioning is the process of precisely placing one end of an optical fiber at the corresponding target position on the focal plane in modern spectroscopic surveys. Fiber positioning ensures that light from celestial objects is efficiently transmitted to the spectrograph, allowing accurate spectral information of the target object to be acquired. Related technologies can achieve fiber coverage of the focal plane by densely arranging multiple positioners on the focal plane. Each positioner has a specific control range, and the fiber positioner's robotic arm rotates to perform fiber positioning. Each fiber transmits light from a single observation target, thereby achieving fiber coverage of the focal plane.
[0003] However, when multiple locators are densely arranged on the focal plane and adjusted, mechanical control is generally performed on each locator separately according to pre-set rule algorithms and paths. The algorithm program running time is long, resulting in low positioning efficiency. On the other hand, the process of adjusting the robotic arm lacks flexibility and adaptability, and the risk of collision is also high. Summary of the Invention
[0004] The main purpose of the embodiments of the present application is to propose a control method, device, computer equipment and storage medium for a fiber optic locator, which can improve the positioning efficiency of the fiber optic locator while effectively reducing the possibility of collision between robotic arms.
[0005] To achieve the above objectives, a first aspect of an embodiment of the present application provides a method for controlling a fiber optic positioner, the method comprising:
[0006] determining a plurality of first optical fiber positioners to be positioned, wherein each of the first optical fiber positioners has at least one rotatable first mechanical arm;
[0007] For each of the first optical fiber positioners, determining at least one second optical fiber positioner adjacent to the first optical fiber positioner and a second robotic arm corresponding to each second optical fiber positioner;
[0008] Obtaining a current first pose of the first robotic arm and a target pose for the first robotic arm;
[0009] Obtain the current second pose of each second robotic arm, and calculate the rotation path based on the first pose of each first robotic arm, the target pose, and the second pose of each adjacent second robotic arm through the target reinforcement learning model to obtain the target rotation path of each first robotic arm to avoid the adjacent second robotic arm;
[0010] Positioning the first robotic arm corresponding to the first optical fiber positioner based on the target rotation route;
[0011] The target reinforcement learning model is obtained by adjusting the model parameters of the target optimization reward constructed by the initial reinforcement learning model based on the sample first pose, the sample first intermediate pose and the sample target pose;
[0012] Among them, the sample first posture is the sample first posture of the sample first robotic arm of the sample first optical fiber locator, the sample first intermediate posture is the posture obtained by adjusting the sample first robotic arm according to the sample target rotation route corresponding to the perceived sample environment information, and the sample target posture is the target posture for the sample first robotic arm; the sample environment information at least includes the posture change information of the sample first intermediate posture of the sample first robotic arm of the sample first optical fiber locator, and the posture change information of the sample second posture of each sample second robotic arm corresponding to at least one sample second optical fiber locator adjacent to the sample first optical fiber locator.
[0013] Accordingly, a second aspect of an embodiment of the present application provides a control device for a fiber optic positioner, the device comprising:
[0014] a first determining module, configured to determine a plurality of first optical fiber positioners to be positioned, wherein each of the first optical fiber positioners has at least one rotatable first mechanical arm;
[0015] a second determining module configured to determine, for each of the first optical fiber positioners, at least one second optical fiber positioner adjacent to the first optical fiber positioner and a second mechanical arm corresponding to each second optical fiber positioner;
[0016] A first acquisition module is used to acquire a current first pose of the first robotic arm and a target pose for the first robotic arm;
[0017] A second acquisition module is used to obtain the current second posture of each second robotic arm, and calculate the rotation path based on the first posture, target posture and the second posture of each adjacent second robotic arm of each first robotic arm through the target reinforcement learning model to obtain the target rotation path of each first robotic arm in avoiding the adjacent second robotic arm;
[0018] A positioning module for positioning the first robotic arm corresponding to the first optical fiber locator based on the target rotation route; wherein the target reinforcement learning model is obtained after the target optimization reward adjustment model parameters are constructed by the initial reinforcement learning model based on the sample first pose, the sample first intermediate pose and the sample target pose; wherein the sample first pose is the sample first pose of the sample first robotic arm of the sample first optical fiber locator, the sample first intermediate pose is the pose obtained by adjusting the sample first robotic arm according to the sample target rotation route corresponding to the perceived sample environment information, and the sample target pose is the target pose for the sample first robotic arm; the sample environment information at least includes the pose change information of the sample first intermediate pose of the sample first robotic arm of the sample first optical fiber locator, and the pose change information of the sample second pose of each sample second robotic arm corresponding to at least one sample second optical fiber locator adjacent to the sample first optical fiber locator.
[0019] In some embodiments, the control device of the optical fiber positioner further includes a training module for:
[0020] obtaining a sample first position of at least one sample first mechanical arm of a sample first optical fiber positioner;
[0021] Perceiving sample environment information corresponding to the first optical fiber locator of the sample through an initial reinforcement learning model, and obtaining a sample target rotation path;
[0022] Adjusting the posture of each sample first manipulator according to the sample target rotation route, and determining the sample first intermediate posture after the adjustment of the sample first manipulator;
[0023] Obtaining a sample target pose for each of the sample first robotic arms, and constructing a target optimization reward based on the sample first pose, the sample first intermediate pose, and the sample target pose;
[0024] The model parameters of the initial reinforcement learning model are adjusted according to the target optimization reward to obtain a target reinforcement learning model.
[0025] In some embodiments, the training module is further configured to:
[0026] At a first sample time, determining a collision reward based on the distance between the sample first fiber locator and a plurality of adjacent sample second fiber locators; wherein the first sample time is the time after the initial reinforcement learning model adjusts the sample first manipulator's sample target rotation path based on the sample environment information perceived;
[0027] determining a target distance reward based on a distance between at least one sample first manipulator of each sample first optical fiber positioner and a sample target posture after the at least one sample first manipulator of each sample first optical fiber positioner rotates from the sample first posture to the sample first intermediate posture;
[0028] Determining a time reward based on the interaction time between the current initial reinforcement learning model and the sample environment information;
[0029] Determining a completion indicator reward based on the positioning status of the plurality of sample first optical fiber locators;
[0030] Determine a target optimization reward of the initial reinforcement learning model at the first sample time based on the collision reward, the target distance reward, the time reward, and the completion index reward.
[0031] In some embodiments, the training module is further configured to:
[0032] At a first sample time, for each of the sample first optical fiber locators, obtaining a minimum sample distance between the profiles of the sample first optical fiber locator and each adjacent sample second optical fiber locator;
[0033] Accumulating multiple minimum sample distances between each of the sample first optical fiber locators and multiple adjacent sample second optical fiber locators to determine a first sample distance;
[0034] For a plurality of sample first manipulators of a first optical fiber positioner of a plurality of samples to be positioned, obtaining a plurality of corresponding first sample distances;
[0035] Based on the multiple first sample distances, a collision reward of the initial reinforcement learning model at the first sample time is determined.
[0036] In some embodiments, the training module is further configured to:
[0037] At a first sample moment, determining a sample first intermediate posture obtained by adjusting each of the sample first mechanical arms of the sample first optical fiber positioner from a sample first posture according to a sample target rotation route corresponding to the sensed sample environment information;
[0038] At the second sample moment, determining a sample second intermediate pose obtained by adjusting each of the sample first manipulators from the sample first intermediate pose according to a sample target rotation route corresponding to the perceived sample environment information;
[0039] Acquire a sample target posture that each of the sample first mechanical arms of the sample first optical fiber positioner needs to position;
[0040] For each of the sample first mechanical arms of each of the sample first optical fiber positioners, obtaining a first difference in rotation angle between the sample first posture and the sample target posture, and obtaining a second difference in rotation angle between the sample second posture and the sample target posture;
[0041] determining a first target distance reward for each of the sample first fiber locators based on the first difference and the second difference of each of the sample first robotic arms;
[0042] For a plurality of sample first optical fiber locators to be located, a plurality of corresponding first target distance rewards are obtained, and based on the plurality of first target distance rewards, a target distance reward of the initial reinforcement learning model is determined.
[0043] In some embodiments, the training module is further configured to:
[0044] The model parameters of the initial reinforcement learning model are adjusted based on the target optimization reward to obtain the motion state of the sample first optical fiber locator; wherein the motion state includes a first state when the sample first optical fiber locator is successfully positioned to the sample target posture, a second state when the sample first optical fiber locator collides with at least one adjacent sample second optical fiber locator after movement, and a third state when the sample first optical fiber locator is not successfully positioned after movement and does not collide with the adjacent sample second optical fiber locator;
[0045] Obtaining a total number of samples of a plurality of sample first optical fiber locators participating in reinforcement learning training, and when a first ratio of a first number of the plurality of sample first optical fiber locators in a first state, characterized by the motion state, to the total number of samples is lower than a first threshold, repeatedly adjusting model parameters of the initial reinforcement learning model according to the target optimization reward to obtain a new motion state of the sample first optical fiber locator;
[0046] Until the first ratio of the first number of the first fiber locators in the first state of the motion state characterizing multiple samples to the total number of samples is higher than the first threshold, the adjustment of the model parameters of the initial reinforcement learning model is stopped to obtain the target reinforcement learning model.
[0047] In some embodiments, the control device of the optical fiber positioner further includes a computing module for:
[0048] Obtaining motion states corresponding to multiple training rounds, where each training round corresponds to an adjustment of model parameters;
[0049] For each of the training rounds, obtaining a second ratio of a second number of the first optical fiber positioners in the first state of the plurality of samples representing the motion state to the total number of the samples;
[0050] Obtaining a second threshold, and comparing the second ratio with the second threshold to obtain a comparison result;
[0051] Allocating the first model parameters of the current training round to the corresponding probability model based on the comparison result, wherein the probability model includes a first probability distribution model and a second probability distribution model;
[0052] calculating an expected improvement value in a corresponding probability model based on the first model parameter, and determining a second model parameter from the corresponding probability model based on the expected improvement value;
[0053] The second model parameters are used as updated model parameters of the initial reinforcement learning model.
[0054] Correspondingly, the third aspect of the embodiments of the present application proposes a computer device, which includes a memory and a processor, the memory stores a computer program, and the processor implements the control method of the optical fiber locator described in any one of the embodiments of the first aspect of the present application when executing the computer program.
[0055] Correspondingly, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the control method of the optical fiber locator described in any one of the embodiments of the first aspect of the present application.
[0056] The embodiment of the present application obtains and determines a plurality of first optical fiber positioners to be positioned, wherein each of the first optical fiber positioners has at least one rotatable first robotic arm; for each of the first optical fiber positioners, determines at least one second optical fiber positioner adjacent to the first optical fiber positioner and a second robotic arm corresponding to each second optical fiber positioner; obtains the current first posture of the first robotic arm and the target posture for the first robotic arm; obtains the current second posture of each second robotic arm, and calculates a rotation route based on the first posture, target posture and the second posture of each adjacent second robotic arm through a target reinforcement learning model, to obtain a target rotation route for each first robotic arm in avoiding the adjacent second robotic arm; and positions the first robotic arm corresponding to the first optical fiber positioner based on the target rotation route. ; wherein the target reinforcement learning model is obtained after the target optimization reward adjustment model parameters are constructed by the initial reinforcement learning model based on the sample first pose, the sample first intermediate pose, and the sample target pose; wherein the sample first pose is the sample first pose of the sample first manipulator of the sample first fiber locator, the sample first intermediate pose is the pose obtained by adjusting the sample first manipulator according to the sample target rotation route corresponding to the perceived sample environment information, and the sample target pose is the target pose for the sample first manipulator; the sample environment information at least includes the pose change information of the sample first intermediate pose of the sample first manipulator of the sample first fiber locator, and the pose change information of the sample second pose of each sample second manipulator corresponding to at least one sample second fiber locator adjacent to the sample first fiber locator. In this way, by simulating the interaction of multiple fiber locators in the environment, the initial reinforcement learning model can be run and trained multiple times under the constantly changing sample environment information, continuously learning and adjusting strategies, thereby being able to learn how to accurately adjust the rotation route of the manipulator in a complex spatial layout to avoid collisions with other fiber locators, and ultimately achieving high-efficiency and low-collision-rate fiber positioning in practical applications. In summary, the present application can improve the positioning efficiency of the optical fiber locator while effectively reducing the possibility of collision between robotic arms. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 Schematic diagram of the application of the control system of the optical fiber positioner provided in the embodiment of the present application in astronomical spectrum observation;
[0058] Figure 2 is a structural diagram of a fiber optic locator provided in an embodiment of the present application;
[0059] Figure 3 Schematic diagram of a simulated telescope focal plane and optical fiber positioner provided in an embodiment of the present application;
[0060] Figure 4 is a flow chart of a control method of an optical fiber locator provided in an embodiment of the present application;
[0061] Figure 5 This is a schematic diagram of the simulated distribution of optical fiber locators provided in an embodiment of the present application;
[0062] Figure 6 Schematic diagram of a reinforcement learning network provided in an embodiment of the present application;
[0063] Figure 7 This is a training flow chart for the initial reinforcement learning model provided in an embodiment of the present application;
[0064] Figure 8 This is a schematic diagram of the functional modules of the control device of the optical fiber positioner provided in an embodiment of the present application;
[0065] Figure 9 This is a schematic diagram of the hardware structure of the computer device provided in the embodiment of the present application. DETAILED DESCRIPTION
[0066] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0067] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.
[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0069] Fiber positioning is the process of precisely placing one end of an optical fiber at the corresponding target position on the focal plane in modern spectroscopic surveys. Fiber positioning ensures that light from celestial objects is efficiently transmitted to the spectrograph, allowing accurate spectral information of the target object to be acquired. Related technologies can achieve fiber coverage of the focal plane by densely arranging multiple positioners on the focal plane. Each positioner has a specific control range, and the fiber positioner's robotic arm rotates to perform fiber positioning. Each fiber transmits light from a single observation target, thereby achieving fiber coverage of the focal plane.
[0070] However, when multiple locators are densely arranged on the focal plane and adjusted, mechanical control is generally performed on each locator separately according to pre-set rule algorithms and paths. The algorithm program running time is long, resulting in low positioning efficiency. In future large-scale spectral surveys, the number of fiber optic locators will reach tens of thousands, and the existing fiber optic positioning method will not be able to meet the new positioning needs. On the other hand, the process of adjusting the robotic arm lacks flexibility and adaptability, and the risk of collision is also high.
[0071] Based on this, the embodiments of the present application provide a control method, device, computer equipment and storage medium for a fiber optic locator, which can improve the positioning efficiency of the fiber optic locator while effectively reducing the possibility of collision between robotic arms.
[0072] The control method, device, computer equipment and storage medium of the fiber optic locator provided in the embodiments of the present application are specifically illustrated through the following embodiments. First, the application of the control system of the fiber optic locator in the embodiments of the present application in astronomical spectral observation is described.
[0073] Please refer to Figure 1 In some embodiments, the light from the celestial body passes through the optical system of the telescope and hits the focal plane. One end of several optical fibers is arranged in the focal plane system, and the other end of the optical fiber is connected to the spectrograph. With this mechanism, the light from the celestial body is transmitted to the spectrograph through the optical fiber, thereby capturing the spectral data of the celestial body. Each optical fiber transmits the light of one observation target. By arranging multiple optical fibers on the focal plane, it is possible to simultaneously transmit the light of multiple observation targets, and then simultaneously capture spectra of multiple observation target celestial bodies, thereby improving the efficiency of the sky survey. Before observation, one end of each optical fiber is placed at the position of its corresponding target on the focal plane to prepare for receiving the light of this target. This placement process is called fiber positioning.
[0074] Each fiber optic locator can have at least one mechanical arm, for example, 1, 2, 3, etc. As long as the fiber optic locator involves adjusting the mechanical arm or components to achieve observation, it is applicable to the technical solution of this application. Figure 2 As shown, the fiber optic positioner resembles a human arm, with two connected arms corresponding to the upper and lower arms, respectively. The outer end of the arm, corresponding to the human hand, controls one end of the optical fiber. This structure allows the fiber optic positioner to position the fiber by controlling the rotation of the two axes. Each fiber optic positioner has a specific control range, defined as a circle with the sum of the lengths of the two arms as its radius. By densely arranging multiple positioners on the focal plane, fiber coverage of the focal plane is achieved.
[0075] In some embodiments, the present application provides a control system for a fiber optic locator. The control system for the fiber optic locator is based on the above-mentioned fiber optic locator and constructs a reinforcement learning network. Specifically, the control system for the fiber optic locator can be the focal plane of an actual telescope and multiple fiber optic locators. The number of fiber optic locators can be 60, 100, 1000, etc. For example, the control system for the fiber optic locator can be distributed with 60 focal planes of fiber optic locators. Figure 3 As shown, Figure 3 To simulate the focal plane and fiber optic locator of the telescope, multiple fiber optic locators can be divided into 6 triangular modules, each module contains 10 fiber optic locators, and the 6 triangular modules are arranged in a hexagonal shape as a whole. The two robotic arms of the fiber optic locator use semi-circular rectangle and round-head trapezoidal outlines respectively.
[0076] Furthermore, the fiber optic positioner's control system can use a reinforcement learning model to conduct multiple rounds of interactions within the environment. During each interaction, it observes the state of the sample environment information, takes actions in response to the sample environment, and then receives rewards from the environment and makes new observations. During this process, the reinforcement learning model can learn to adjust the robot arm's angle to avoid collisions, ultimately achieving precise positioning of the robot arm.
[0077] The control method of the optical fiber positioner in the embodiment of the present application can be illustrated by the following embodiment.
[0078] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0079] In the embodiment of the present application, the control device of the optical fiber positioner will be described from the perspective of the control device of the optical fiber positioner, which can be integrated into a computer device. Figure 4 , Figure 4 This is a flowchart of the steps of the control method of the optical fiber locator provided in an embodiment of the present application. In this embodiment of the present application, the control device of the optical fiber locator is specifically integrated into a terminal or server as an example. When the processor on the terminal or server executes the program instructions corresponding to the control method of the optical fiber locator, the specific process is as follows:
[0080] Step 101: Determine a plurality of first optical fiber positioners to be positioned, wherein each first optical fiber positioner has at least one rotatable first mechanical arm.
[0081] In some embodiments, to improve positioning accuracy and efficiency, multiple first optical fiber locators may be acquired to facilitate parallel processing and distributed positioning of the first optical fiber locators, thereby improving positioning speed and accuracy.
[0082] The first fiber optic positioner may be a fiber optic positioner that needs to adjust its rotation path to accurately position itself at the target position. After the first fiber optic positioner is adjusted to the target position, the optical fiber can be positioned, thereby directing the optical signal onto the focal plane of the telescope.
[0083] The first robotic arm can be part of a first fiber locator, connected to the optical fiber and capable of positioning the optical fiber through rotation. The rotation angle of the first robotic arm can be adjusted using a target reinforcement learning model to achieve optimal positioning. For example, each first fiber locator can have one, two, or more first robotic arms, and the number of first robotic arms in each first fiber locator can be the same or different.
[0084] For example, all fiber locators can be scanned by sensors in the system to determine the current status of all fiber locators, so as to determine whether there are fiber locators among the multiple fiber locators that have been positioned, are damaged, or do not need to be positioned, and these fiber locators can be screened to obtain the first fiber locator that needs to be positioned.
[0085] Exemplarily, the optical fiber positioners may not be screened, and all optical fiber positioners may be used as the first optical fiber positioners.
[0086] Through the above method, the first optical fiber positioner that needs to perform rotation route calculation can be determined, so as to facilitate the subsequent positioning of the first mechanical arm of the optical fiber positioner.
[0087] Step 102: For each first fiber locator, determine at least one second fiber locator adjacent to the first fiber locator and a second robotic arm corresponding to each second fiber locator.
[0088] In some embodiments, in order to provide a valid state representation for the target reinforcement learning model so that the target reinforcement learning model can make the best action decision and avoid collisions between the first fiber locators during the positioning process, at least one adjacent second fiber locator and the second robotic arm corresponding to the second fiber locator can be determined for each first fiber locator, so that the target reinforcement learning model can optimize the action strategy according to the second robotic arm corresponding to the adjacent second fiber locator to achieve efficient and high-precision fiber positioning.
[0089] The second optical fiber locator may be an optical fiber locator that is adjacent to the first optical fiber locator and has a collision probability with the first optical fiber locator.
[0090] The second robotic arm can be part of a second fiber locator, connected to the optical fiber and capable of positioning the optical fiber through rotation. The rotation angle of the second robotic arm can be adjusted using a target reinforcement learning model to achieve optimal positioning. For example, each second fiber locator can have one, two, or more second robotic arms, and the number of second robotic arms in each second fiber locator can be the same or different.
[0091] For example, please refer to Figure 5 ,exist Figure 5 In the example, each fiber optic positioner can be a first fiber optic positioner. Figure 5 All the optical fiber locators in the are adjusted and positioned, then the optical fiber locators a, b, c, d, e, f, g, h, i can all be the first optical fiber locators. For the first optical fiber locator e, the optical fiber locators a, b, c, d, f, g, h, i are all second optical fiber locators. For the first optical fiber locator i, the optical fiber locators e, f, h are all second optical fiber locators.
[0092] Through the above method, the adjacent second fiber optic locator can be determined for each first fiber optic locator, thereby avoiding collision between the fiber optic locators during the positioning process. As a result, the target reinforcement learning model can optimize the rotation angle of the first robotic arm according to the state of the adjacent second fiber optic locator and its second robotic arm, thereby achieving efficient and high-precision fiber optic positioning and improving the positioning success rate and reliability of the overall system.
[0093] Step 103: Obtain the current first pose of the first robotic arm and the target pose for the first robotic arm.
[0094] In some embodiments, in order to facilitate the calculation of the rotation route of the first robotic arm of the first fiber optic positioner, the current first posture of the first robotic arm and the target posture to which the first robotic arm ultimately needs to be positioned can be obtained, so as to plan the rotation route of the first robotic arm according to the target posture.
[0095] The first pose may be the current position and attitude of the first robotic arm corresponding to the first fiber locator. The first pose includes the position (e.g., x, y, z coordinates) and attitude (e.g., pitch, roll, and yaw angle information) of the first robotic arm.
[0096] The target posture may be the posture that the first robotic arm corresponding to the first fiber optic positioner ultimately needs to achieve (e.g., pitch, roll, and yaw angle information, etc.). Furthermore, the target posture may also include the position that the first robotic arm ultimately needs to achieve (e.g., x, y, and z coordinates). The target posture may be pre-set based on the task requirements of the fiber optic positioning.
[0097] For example, assuming that the first fiber optic positioner a needs to adjust the position of the first robotic arm, then the current position and posture of the first robotic arm corresponding to a is the first posture, and the final position and posture to which the first robotic arm corresponding to a needs to be adjusted is the target posture.
[0098] By obtaining the current first posture of the first robotic arm and the target posture of the first robotic arm, it is possible to facilitate the subsequent calculation of the rotation route required for the first robotic arm of the first fiber optic positioner to reach the target posture, thereby optimizing the motion path of the first robotic arm and reducing uncertainty and possible collision risks during the motion process.
[0099] Step 104: Obtain the current second posture of each second robotic arm, and calculate the rotation path based on the first posture, target posture and the second posture of each adjacent second robotic arm through the target reinforcement learning model to obtain the target rotation path of each first robotic arm in avoiding the adjacent second robotic arm.
[0100] In some embodiments, in order to effectively respond to a changing environment, the current second posture of each second robotic arm can be obtained, and a trained target reinforcement learning model can be used to calculate the optimal rotation route for the first robotic arm to reach the target posture while avoiding the adjacent second robotic arm based on the current posture of the first robotic arm, the target posture, and the current posture of the adjacent second robotic arm, so as to effectively respond to changes in the environment and achieve efficient, safe, and accurate fiber optic positioning operations.
[0101] Among them, the second posture can be the position (such as x, y, and z coordinates) and posture state (such as pitch angle, roll angle, yaw angle, etc.) of the second robotic arm corresponding to the second fiber optic locator at the same time when the first robotic arm of the first fiber optic locator is in the first posture.
[0102] Among them, the target reinforcement learning model can be a trained intelligent agent, and the target reinforcement learning model can automatically, quickly and accurately plan the target rotation route through the first posture change of the first robotic arm of the first fiber optic positioner and the second posture change of the second robotic arm of the second fiber optic positioner.
[0103] Among them, the target rotation route can be a series of rotation action routes that the first robotic arm needs to perform in order to reach the target posture. The target rotation route can include the angle at which the first robotic arm needs to rotate, and can also include the order of rotation, etc., to ensure that the first robotic arm can move smoothly and accurately to the target posture.
[0104] For example, for the first fiber locator A, the adjacent second fiber locators B, C, and D, as well as the second robotic arm corresponding to the second fiber locators B, C, and D, can be determined, and the current position of the second robotic arm is the second position. Furthermore, the first position, the target position, and the second position can be input into a target reinforcement learning model. The target reinforcement learning model can calculate the target rotation path of the first fiber locator based on the first position, the target position, and the second position, and obtain the optimal rotation path for the first robotic arm to reach the target position, that is, the target rotation path.
[0105] Through the above method, the target rotation route can be obtained, so that the first robotic arm of the subsequent first optical fiber positioner can accurately reach the target position according to the target rotation route, and also avoid collision with the adjacent second robotic arm, thereby improving the efficiency and safety of optical fiber positioning.
[0106] Step 105 : Positioning the first mechanical arm corresponding to the first optical fiber positioner based on the target rotation path.
[0107] The target reinforcement learning model is obtained by adjusting the model parameters by the target optimization reward constructed by the initial reinforcement learning model based on the sample first pose, the sample first intermediate pose, and the sample target pose;
[0108] Among them, the sample first posture is the sample first posture of the sample first robotic arm of the sample first fiber locator, the sample first intermediate posture is the posture obtained by adjusting the sample first robotic arm according to the sample target rotation route corresponding to the perceived sample environment information, and the sample target posture is the target posture for the sample first robotic arm; the sample environment information at least includes the posture change information of the sample first intermediate posture of the sample first robotic arm of the sample first fiber locator, and the posture change information of the sample second posture of each sample second robotic arm corresponding to at least one sample second fiber locator adjacent to the sample first fiber locator.
[0109] In some embodiments, in order to achieve precise position and posture control of the first fiber optic positioner, the first robotic arm corresponding to the first fiber optic positioner can be positioned based on the target rotation route to ensure that the first robotic arm of the first fiber optic positioner can smoothly reach the predetermined working position while avoiding collision with other second robotic arms or obstacles.
[0110] The initial reinforcement learning model may be an untrained reinforcement learning model. By training the initial reinforcement learning model and continuously adjusting the model parameters, the target reinforcement learning model may be obtained.
[0111] The sample first pose may be the unadjusted position and attitude state of the sample first manipulator corresponding to the sample first fiber locator during the training of the initial reinforcement learning model. The sample first pose includes the position (e.g., x, y, z coordinates) and attitude (e.g., pitch, roll, and yaw angle information, etc.) of the sample first manipulator.
[0112] The sample first intermediate pose may be the current position and attitude of the sample first manipulator corresponding to the sample first fiber locator after at least one adjustment during the training of the initial reinforcement learning model. The sample first pose includes the position (e.g., x, y, z coordinates) and attitude (e.g., pitch, roll, and yaw angle information, etc.) of the sample first manipulator.
[0113] The sample target pose can be the pose that the first robotic arm corresponding to the first fiber optic positioner ultimately needs to achieve (such as pitch, roll, and yaw angle information). Furthermore, the sample target pose can also include the position that the first robotic arm ultimately needs to achieve (such as x, y, and z coordinates). The sample target pose can be pre-set based on the task requirements of the fiber optic positioning.
[0114] Among them, the target optimization reward can be used to guide the initial reinforcement learning model to learn towards the sample target pose corresponding to the sample first robotic arm, and guide the initial reinforcement learning model to learn how to make optimal decisions in the process of interacting with the sample environment.
[0115] The sample first optical fiber positioner may be a sample optical fiber positioner that needs to adjust its rotation path to accurately position itself at the target position. After adjusting the sample first optical fiber positioner to the sample target position, the optical fiber can be positioned, thereby directing the optical signal to the focal plane of the telescope.
[0116] The sample first robotic arm can be part of a sample first fiber optic positioner, which can be connected to the optical fiber and achieve optical fiber positioning through rotation. The rotation angle of the sample first robotic arm can be adjusted by the initial reinforcement learning model to achieve better positioning effect. Exemplarily, each sample first fiber optic positioner can have one, two, or more sample first robotic arms, and the number of sample first robotic arms of each sample first fiber optic positioner can be the same or different.
[0117] The sample environment information may be data used by the initial reinforcement learning model to perceive and understand the environment during the reinforcement learning process of the initial reinforcement learning model. For example, the sample environment information may include position change information of a sample first intermediate pose of a sample first manipulator of a sample first fiber locator, and position change information of a sample second pose of each sample second manipulator corresponding to at least one sample second fiber locator adjacent to the sample first fiber locator.
[0118] Among them, the sample target rotation route can be a series of rotation action routes that the first robotic arm needs to perform in order to reach the sample target posture. The sample target rotation route can include the angle at which the sample first robotic arm needs to rotate, and can also include the order of rotation, etc., to ensure that the sample first robotic arm can move smoothly and accurately to the sample target posture.
[0119] The sample second optical fiber locator may be a sample optical fiber locator that is adjacent to the sample first optical fiber locator and has a collision probability.
[0120] The sample second robotic arm can be part of a sample second fiber optic positioner, can be connected to an optical fiber, and can achieve optical fiber positioning through rotation. The rotation angle of the sample second robotic arm can be adjusted by an initial reinforcement learning model to achieve optimal positioning. Exemplarily, each sample second fiber optic positioner can have one, two, or more sample second robotic arms, and the number of sample second robotic arms of each sample second fiber optic positioner can be the same or different.
[0121] Among them, the sample second posture can be the position (such as x, y, and z coordinates) and posture state (such as pitch angle, roll angle, yaw angle, etc.) of the sample second robotic arm corresponding to the sample second optical fiber locator at the same moment when the sample first robotic arm of the sample first optical fiber locator is in the sample first posture.
[0122] For example, after obtaining the target rotational path, the first robotic arm corresponding to the first fiber positioner can be controlled to translate or rotate. Furthermore, a fiber grating sensor can be installed on the first robotic arm to monitor its position and motion. By using this data for real-time feedback, the rotation and movement of the first robotic arm can be precisely adjusted to ensure that it is accurately adjusted to the target position according to the predetermined target rotational path.
[0123] For example, after obtaining the target rotation path, the first robotic arm can be controlled by a magnetic field, or the rotation of the first robotic arm can be controlled in combination with a visual system, etc. This application does not limit the specific method of rotating to the target posture based on the target rotation path.
[0124] For example, the reinforcement learning network constructed can be a soft actor-critic (SAC) reinforcement learning network. Figure 6 As shown in Figure 1, SAC is an off-policy algorithm. The initial reinforcement learning model can interact with the environment multiple times and store the interaction data in a buffer. After a certain number of interactions, a random portion of the interactions are read from the buffer for training. SAC incorporates techniques such as entropy, target Q-networks and their soft updates, and dual-critic networks to optimize training efficiency and mitigate overfitting. The neural network in SAC consists of an actor network, two Q-networks, and two target Q-networks. These networks can be implemented using multilayer perceptrons (MLPs).
[0125] In some embodiments, when the initial reinforcement learning model begins to be trained, the sample target pose can randomly generate a sample target rotation route, and then the target optimization reward is calculated by controlling the sample first robotic arm. The initial reinforcement learning model can adjust the model parameters according to the target optimization reward and regenerate a new sample target rotation route. The initial reinforcement learning model is repeatedly trained to finally obtain the target reinforcement learning model.
[0126] It can be understood that during the reinforcement learning process of the initial reinforcement learning model, the sample environment information is also constantly changing. After sensing the environmental changes, the initial reinforcement learning model can generate a new sample target rotation route and position the sample first robotic arm corresponding to the sample first fiber optic locator based on the new sample target rotation route.
[0127] Furthermore, after the number of reinforcement learning times of the initial reinforcement learning model exceeds a preset number, for example, after more than 500 times, the training of the initial reinforcement learning model can be stopped to obtain the target reinforcement learning model. In another embodiment, after the initial reinforcement learning model successfully locates multiple first fiber locators to the sample target posture, the number of successfully located first fiber locators can be counted. When the ratio of the number of successfully located first fiber locators to the total number of first fiber locators exceeds a first threshold, such as the first threshold is 80%, when the ratio of the number of successfully located first fiber locators to the total number of first fiber locators is 85%, the adjustment of the model parameters of the initial reinforcement learning model can be stopped to obtain the target reinforcement learning model.
[0128] The embodiment of the present application obtains and determines a plurality of first optical fiber positioners to be positioned, wherein each of the first optical fiber positioners has at least one rotatable first robotic arm; for each of the first optical fiber positioners, determines at least one second optical fiber positioner adjacent to the first optical fiber positioner and a second robotic arm corresponding to each second optical fiber positioner; obtains the current first posture of the first robotic arm and the target posture for the first robotic arm; obtains the current second posture of each second robotic arm, and calculates a rotation route based on the first posture, target posture and the second posture of each adjacent second robotic arm through a target reinforcement learning model, to obtain a target rotation route for each first robotic arm in avoiding the adjacent second robotic arm; and positions the first robotic arm corresponding to the first optical fiber positioner based on the target rotation route. ; wherein the target reinforcement learning model is obtained after the target optimization reward adjustment model parameters are constructed by the initial reinforcement learning model based on the sample first pose, the sample first intermediate pose, and the sample target pose; wherein the sample first pose is the sample first pose of the sample first manipulator of the sample first fiber locator, the sample first intermediate pose is the pose obtained by adjusting the sample first manipulator according to the sample target rotation route corresponding to the perceived sample environment information, and the sample target pose is the target pose for the sample first manipulator; the sample environment information at least includes the pose change information of the sample first intermediate pose of the sample first manipulator of the sample first fiber locator, and the pose change information of the sample second pose of each sample second manipulator corresponding to at least one sample second fiber locator adjacent to the sample first fiber locator. In this way, by simulating the interaction of multiple fiber locators in the environment, the initial reinforcement learning model can be run and trained multiple times under the constantly changing sample environment information, continuously learning and adjusting strategies, thereby being able to learn how to accurately adjust the rotation route of the manipulator in a complex spatial layout to avoid collisions with other fiber locators, and ultimately achieving high-efficiency and low-collision-rate fiber positioning in practical applications. In summary, the present application can improve the positioning efficiency of the optical fiber locator while effectively reducing the possibility of collision between robotic arms.
[0129] Please refer to Figure 7In some embodiments, in order to enable the initial reinforcement learning model to learn how to effectively perform tasks in a given environment, the initial reinforcement learning model can be trained so that the first fiber locator can complete the task efficiently and accurately. Furthermore, when training the initial reinforcement learning model, it can be trained through actual scenarios or through simulated scenarios. When simulating the scenario, the computer can simulate the focal plane of the telescope and the fiber locator. The simulated scene includes the center coordinates of each fiber locator, the contour shape of the two mechanical arms of the fiber locator, the rotation angle of each mechanical arm, and the position and contour shape of other fiber locator instruments on the focal plane that need to be considered for collision with the fiber locator. For example, the target reinforcement learning model used in the present application mentioned above is trained in the following way:
[0130] Step 201: Acquire a sample first pose of at least one sample first mechanical arm of a sample first optical fiber positioner.
[0131] In some embodiments, to improve positioning accuracy and efficiency, multiple sample first optical fiber locators may be acquired to facilitate parallel processing and distributed positioning of the sample first optical fiber locators, thereby improving positioning speed and accuracy.
[0132] For example, all fiber locators can be scanned by sensors in the system to determine the current status of all fiber locators, so as to determine whether there are fiber locators among the multiple fiber locators that have been positioned, are damaged, or do not need to be positioned, and these fiber locators can be screened to obtain a sample first fiber locator that needs to be positioned.
[0133] For example, the optical fiber positioners may not be screened, and all optical fiber positioners may be used as sample first optical fiber positioners.
[0134] Through the above method, the sample first optical fiber positioner that needs to be rotated can be determined, so as to facilitate the subsequent positioning of the sample first mechanical arm of the optical fiber positioner.
[0135] Step 202 : The sample environment information corresponding to the sample first optical fiber locator is perceived by the initial reinforcement learning model to obtain the sample target rotation path.
[0136] In some embodiments, in some embodiments, in order to effectively respond to a changing environment, the current sample second posture of each sample second robotic arm can be obtained, and an initial reinforcement learning model can be used to calculate the rotation route of the sample first robotic arm to reach the sample target posture while avoiding the adjacent sample second robotic arm based on the sample first posture of the sample first robotic arm, the sample target posture, and the second posture of the adjacent sample second robotic arm, so as to effectively respond to changes in the environment.
[0137] Exemplarily, the initial reinforcement learning model can calculate the sample target rotation route based on the sample first posture of the sample first robotic arm of the sample first fiber locator and the sample second posture of each sample second robotic arm corresponding to at least one sample second fiber locator adjacent to the sample first fiber locator.
[0138] Through the above method, the sample target rotation route can be obtained, so that the sample first robot arm of the subsequent sample first fiber optic positioner can accurately reach the sample target position according to the sample target rotation route, and also avoid collision with the adjacent sample second robot arm, thereby improving the efficiency and safety of fiber optic positioning.
[0139] Step 203 : Adjust the posture of each sample first manipulator according to the sample target rotation route, and determine the sample first intermediate posture after the adjustment of the sample first manipulator.
[0140] In some embodiments, in order to achieve precise position and posture control of the sample first fiber optic locator, the sample first robotic arm corresponding to the sample first fiber optic locator can be adjusted based on the sample target rotation route to ensure that the sample first robotic arm of the sample first fiber optic locator can smoothly reach the predetermined working position while avoiding collision with other sample second robotic arms or obstacles.
[0141] For example, since the absolute value of the rotation angle change of the sample robotic arm of each sample fiber optic positioner has an upper limit constraint, such as 5°, when a large adjustment amplitude is required, the sample first robotic arm may not immediately reach the sample target posture. Therefore, the first intermediate posture can be an intermediate posture in the sample target rotation route, or it can be the sample target posture.
[0142] For example, after obtaining the sample target rotation path, the initial reinforcement learning model can control the translation or rotation of the sample first manipulator corresponding to the sample first fiber optic positioner. Furthermore, a fiber grating sensor can be installed on the sample first manipulator to monitor its position and motion. By using this data for real-time feedback, the rotation and movement of the sample first manipulator can be precisely adjusted to ensure that the sample first manipulator is accurately adjusted to the sample target position according to the predetermined sample target rotation path.
[0143] For example, after obtaining the sample target rotation route, the sample first robotic arm can be controlled by a magnetic field, and the rotation of the sample first robotic arm can be controlled in combination with a visual system, etc. This application does not limit the specific method of rotating to the sample target posture based on the sample target rotation route.
[0144] By determining the first intermediate pose of the sample after the adjustment of the first manipulator of the sample, it is convenient to subsequently determine the target optimization reward.
[0145] Step 204 : Obtain a sample target pose for each sample first robotic arm, and construct a target optimization reward based on the sample first pose, the sample first intermediate pose, and the sample target pose.
[0146] In some embodiments, in order to guide the initial reinforcement learning model to learn towards the sample target pose corresponding to the sample first robotic arm, a target optimization reward can be constructed to guide the initial reinforcement learning model to learn how to make optimal decisions in the process of interacting with the sample environment.
[0147] The target optimization reward can be a combination of a collision reward, a target distance reward, a time reward, and a completion indicator reward. It can be used to guide the initial reinforcement learning model toward the sample target pose corresponding to the sample first robotic arm and guide the initial reinforcement learning model to learn how to make optimal decisions during interaction with the sample environment.
[0148] Through the above method, a target optimization reward can be constructed to facilitate the subsequent training of the initial reinforcement learning model based on the target optimization reward, so that the initial reinforcement learning model has better decision-making ability.
[0149] In some embodiments, in order to guide the initial reinforcement learning model to learn towards the sample target pose corresponding to the sample first robotic arm, a target optimization reward may be constructed to guide the initial reinforcement learning model to learn how to make optimal decisions during the interaction with the sample environment. For example, the "constructing a target optimization reward based on the sample first pose, the sample first intermediate pose, and the sample target pose" in step 204 may include:
[0150] (204.1) Determine a collision reward at a first sample time based on the distance between the sample first fiber optic locator and multiple adjacent sample second fiber optic locators; wherein the first sample time is the time after the initial reinforcement learning model adjusts the sample first manipulator's sample target rotation path based on the perceived sample environment information;
[0151] (204.2) Determine a target distance reward based on the distance between at least one sample first manipulator of each sample first fiber positioner and the sample target position after the manipulator rotates from the sample first position to the sample first intermediate position;
[0152] (204.3) Determine the time reward based on the interaction time between the current initial reinforcement learning model and the sample environment information;
[0153] (204.4) Determining a completion indicator reward based on the positioning status of the first optical fiber locators of the plurality of samples;
[0154] (204.5) Based on the collision reward, target distance reward, time reward and completion indicator reward, determine the target optimization reward of the initial reinforcement learning model at the first sample moment.
[0155] The first sample moment may be the moment when the initial reinforcement learning model adjusts the sample first fiber locator from the sample first posture to the sample first intermediate posture. In other words, the first sample moment may be the moment when the initial reinforcement learning model adjusts the sample first manipulator to the sample target rotation path.
[0156] The collision reward can be the minimum distance reward between the first manipulator of the sample first fiber positioner and the outlines of multiple adjacent second manipulators after the first manipulator of the sample first fiber positioner moves to the first intermediate pose in the focal plane device. The collision reward is used to incentivize the initial reinforcement learning model to avoid executing actions that may cause collisions between the sample fiber positioners. Specifically, the collision reward can be the minimum distance between the outlines of the first and second manipulators. A distance of 0 indicates contact, while a negative value indicates a collision.
[0157] The target distance reward can be a reward given by the initial reinforcement learning model based on the distance to the sample target pose when the initial reinforcement learning model adjusts the rotation angle of the sample first manipulator of the sample first fiber locator to approach the sample target pose. The target distance reward can be used to motivate the initial reinforcement learning model to approach the sample target pose more quickly. The value of the target distance reward can be the negative value of the difference between the target rotation angle of the sample first manipulator of the sample first positioner and the target angle of the sample target pose. The smaller the difference, the closer the sample target pose is to the sample target pose, and the larger the target distance reward.
[0158] The time reward can be calculated by adjusting the time taken by the initial reinforcement learning model to adjust the first robot arm of the sample for each interaction during the localization task. Therefore, a certain reward is deducted from each interaction step. The time reward can be used to incentivize the initial reinforcement learning model to complete the localization task faster, reducing the required time. The time reward is a negative number. Each additional interaction step by the initial reinforcement learning model consumes more time, so a certain reward is deducted from each interaction step.
[0159] The completion indicator reward can be a reward for successful or failed positioning. For each interaction, if the initial reinforcement learning model successfully adjusts the sample first manipulator of the sample first fiber locator to the sample target pose, a positive reward is given; if a collision causes positioning failure, a negative reward is given; if there is neither success nor failure, the reward is zero. The completion indicator reward is used to incentivize the initial reinforcement learning model to successfully complete the positioning task as quickly as possible and avoid failure.
[0160] Specifically, each time the initial reinforcement learning model is adjusted according to the sample target rotation route, it is necessary to calculate the target optimization reward and adjust the model parameters of the initial reinforcement model according to the target optimization reward at the current moment. After the adjustment, a new sample target rotation route is formulated according to the updated environment, and this is repeated.
[0161] For example, during the training of the initial reinforcement learning model, the sample environment information and the reinforcement learning model can be initialized. The task of the initial reinforcement learning model is to learn how to control the sample first robotic arm of the sample first fiber optic positioner in the sample environment to achieve the sample target posture.
[0162] Furthermore, the initial reinforcement learning model begins to perceive the sample environment information, and based on the perceived sample environment information, adjusts the sample target rotation route of the sample first robotic arm, and calculates the distance between the sample first robotic arm and the adjacent sample second fiber optic locator according to the adjusted sample target rotation route to determine the collision reward.
[0163] Furthermore, after the sample first manipulator rotates from the sample first pose to the sample intermediate pose, the distance between the sample intermediate pose and the sample target pose is calculated to determine the target distance reward.
[0164] Furthermore, a time reward can be determined based on the interaction time between the initial reinforcement learning model and the sample environment information. If the model completes the task in a shorter time, the time reward is higher; if the time is longer, the time reward is lower.
[0165] Furthermore, a reward for achieving a target goal can be determined based on the positioning status of multiple sample first fiber locators. If all sample first fiber locators achieve the expected positioning status, the reward is higher; if any sample first fiber locators do not achieve the expected positioning status, the reward is lower.
[0166] Finally, the collision reward, target distance reward, time reward, and completion indicator reward are added together to obtain the target optimization reward. For example, the target optimization reward can be represented by r, and the calculation formula is as follows:
[0167] r=r c +rd +r t +r s
[0168] Among them, r c is the collision reward, r d is the target distance reward, r t is the time reward, r s Rewards for completing targets.
[0169] Through the above method, the target optimization reward can be calculated. The target optimization reward can be used as a feedback signal for the initial reinforcement learning model to guide the further learning of the initial reinforcement learning model, so that it can learn how to effectively adjust the rotation angle of the robot arm in the fiber optic positioner to achieve the purpose of quickly and accurately positioning the target posture.
[0170] In some embodiments, in order to motivate the initial reinforcement learning model to avoid performing actions that may cause collisions between sample fiber locators, a collision reward for the initial reinforcement learning model may be determined so that the initial reinforcement learning model avoids collisions during subsequent training and application. For example, (204.1) may include:
[0171] (204.1.1) At a first sample time, for each sample first fiber locator, obtain a minimum sample distance between the profiles of the sample first fiber locator and each adjacent sample second fiber locator;
[0172] (204.1.2) Accumulating multiple minimum sample distances between each sample first fiber locator and multiple adjacent sample second fiber locators to determine a first sample distance;
[0173] (204.1.3) Obtaining corresponding first sample distances for a plurality of first sample manipulators of a plurality of first fiber optic positioners of the plurality of samples to be positioned;
[0174] (204.1.4) Based on the multiple first sample distances, determine the collision reward of the initial reinforcement learning model at the first sample time.
[0175] Among them, the minimum sample distance can be the minimum sample distance between the contours of the sample first fiber locator and each adjacent sample second fiber locator calculated at the first sample moment after each sample first manipulator of the sample first fiber locator is adjusted from the sample first posture to the sample first intermediate posture according to the target rotation route based on the perceived sample environment information.
[0176] The first sample distance may be the sum of the minimum sample distances between each sample first optical fiber positioner and all adjacent sample second optical fiber positioners at the first sample time.
[0177] For example, the distance between the sample first fiber locator at the a1 contour point and the sample second fiber locator at the b1 contour point is 0.001 mm, and the minimum sample distance between the sample first fiber locator at the a2 contour point and the sample second fiber locator at the b1 contour point is 0.002 mm. Then, the minimum sample distance between the contours of the sample first fiber locator and each adjacent sample second fiber locator is 0.001 mm.
[0178] Exemplarily, the sample second fiber locators adjacent to the sample first fiber locator include sample second fiber locator A, sample second fiber locator B, and sample second fiber locator C. At the first sample time, the minimum sample distance between the sample first fiber locator and A is 0.01 mm, the minimum sample distance between the sample first fiber locator and B is 0.1 mm, and the minimum sample distance between the sample first fiber locator and C is 0.02 mm. Therefore, the first sample distance is 0.01 + 0.1 + 0.02 = 0.13 mm.
[0179] Furthermore, the first sample distances corresponding to all the first fiber locators of the multiple samples to be positioned can be added together to obtain a collision reward. For example, the sample collision reward r c It can be calculated as follows:
[0180]
[0181] Among them, d ij is the minimum sample distance between the outlines of the sample first fiber locator i and the adjacent sample second fiber locator j. If the minimum sample distance is 0, the sample first fiber locator and the sample second fiber locator are just in contact; if it is a negative value, it indicates that the sample first fiber locator and the sample second fiber locator have collided and positioning failed. p is the total number of sample first fiber locators, n i is the number of sample second fiber locators adjacent to the sample first fiber locator i.
[0182] It is understandable that by squaring the minimum sample distance, the convergence speed of the initial reinforcement learning model can be accelerated.
[0183] The collision reward of the initial reinforcement learning model is calculated in the above manner, which encourages the initial reinforcement learning model to avoid executing actions that may cause collisions between fiber locators, thereby improving the practicality and reliability of the initial reinforcement learning model.
[0184] In some embodiments, in order to encourage the initial reinforcement learning model to approach the sample target pose more quickly, a target distance reward may be calculated to accelerate the convergence of the initial reinforcement learning model. For example, (204.2) may include:
[0185] (204.2.1) at a first sample time, determining a sample first intermediate pose obtained by adjusting each sample first manipulator of the sample first fiber optic positioner from the sample first pose according to a sample target rotation path corresponding to the sensed sample environment information;
[0186] (204.2.2) at the second sample time, determining a second intermediate pose of each sample first manipulator adjusted from the sample first intermediate pose according to a sample target rotation path corresponding to the perceived sample environment information;
[0187] (204.2.3) Obtain the target pose of each sample first manipulator of the sample first fiber optic positioner;
[0188] (204.2.4) For each sample first manipulator of each sample first fiber positioner, obtain a first difference in rotation angle between the sample first pose and the sample target pose, and obtain a second difference in rotation angle between the sample second pose and the sample target pose;
[0189] (204.2.5) Determine a first target distance reward for the first fiber locator of each sample based on the first difference and the second difference of the first robot arm of each sample;
[0190] (204.2.6) For multiple sample first fiber locators to be located, obtain corresponding multiple first target distance rewards, and determine the target distance reward of the initial reinforcement learning model based on the multiple first target distance rewards.
[0191] The second sample moment may be a moment when, during the training process of the initial reinforcement learning model, the sample environment information of the first sample moment changes after the initial reinforcement learning model performs an action on the sample environment information, generating a new state.
[0192] The sample second intermediate pose can be the new pose of the sample first manipulator of the sample first fiber positioner after the initial reinforcement learning model adjusts the sample first manipulator from the sample first intermediate pose. The sample second intermediate pose can be a temporary pose during the rotation process or a sample target pose, depending on the specific situation.
[0193] The first difference may be a rotation angle difference between each sample first mechanical arm of the sample first optical fiber positioner and the sample target posture after the arm rotates from the sample first posture to the sample first intermediate posture.
[0194] The second difference may be a rotation angle difference between each sample first mechanical arm of the sample first optical fiber positioner and the sample target posture after the arm continues to rotate from the sample first intermediate posture to the sample second intermediate posture.
[0195] The first target distance reward may be the difference between the second difference and the first difference, that is, the change in the distance between the sample first robotic arm and the sample target pose after the initial reinforcement learning model takes action.
[0196] For example, when the sample first fiber positioner has two sample first manipulators, the target distance reward r d It can be calculated by the following formula:
[0197]
[0198] Among them, w d is the weight of the target distance reward, p is the total number of sample first fiber locators, α i,t is the first difference, α i,t+1 is the second difference.
[0199] Exemplarily, the calculation formula of the first difference is as follows:
[0200] α i,t =-[(θ d,i -θ i,t ) 2 +(φ d,i -φ i,t ) 2 ]
[0201] Among them, θ d,i ,φ d,i are the target rotation angles of the two sample first manipulators of the sample first fiber positioner i, θ i,t ,φ i,t is the rotation angle of the first manipulator of the two samples of locator i at the first sample time t.
[0202] Exemplarily, the calculation formula for the second difference is as follows:
[0203] α i,t+1 =-[(θ d,i -θ i,t+1 ) 2 +(φ d,i -φ i,t+1 ) 2 ]
[0204] Among them, θ d,i ,φ d,i are the target rotation angles of the two sample first manipulators of the sample first fiber positioner i, θ i,t+1 ,φ i,t+1 is the rotation angle of the first manipulator of the two samples of locator i at the second sample time t.
[0205] In some implementations, a plurality of first target distance rewards corresponding to a plurality of sample first optical fiber locators to be positioned may be added together to obtain a target distance reward for determining an initial reinforcement learning model.
[0206] By calculating the target distance reward of the initial reinforcement learning model, the initial reinforcement learning model can obtain more accurate and timely feedback during the training process, and encourage the initial reinforcement learning model to learn strategies that approach the sample target pose more quickly, thereby accelerating the convergence of the initial reinforcement learning model.
[0207] Step 205 : Adjust the model parameters of the initial reinforcement learning model according to the target optimization reward to obtain the target reinforcement learning model.
[0208] In some embodiments, in order to enable the initial reinforcement learning model to more accurately and efficiently adjust the posture of the sample first robotic arm when facing the fiber optic positioning task, so that the final target reinforcement learning model can show good performance on the given task, the model parameters of the initial reinforcement learning model can be adjusted according to the target optimization reward to obtain the target reinforcement learning model, so as to continuously improve the accuracy of model positioning.
[0209] For example, optimization algorithms such as gradient descent and stochastic gradient descent can be used to adjust model parameters. Specifically, the gradients of the model parameters can be calculated based on the target optimization reward, and the calculated gradients can be used to update the model parameters until the initial reinforcement learning model converges or meets a preset condition, such as reaching a preset number of training times, at which point training is stopped to obtain the target reinforcement learning model.
[0210] Through the above methods, we can adjust the model parameters to maximize the target optimization reward and continuously improve the model performance, so that the target reinforcement learning model has higher adjustment accuracy and efficiency, and reduces the risk of collision.
[0211] In some embodiments, the initial reinforcement learning model can be run for multiple rounds, with each round stopping upon success or failure. Success means that all sample first fiber positioners successfully reach the target rotation angle, and failure means that at least one sample first fiber positioner has collided. The running data is then used to optimize the initial reinforcement learning model until all or most sample first fiber positioners can successfully reach the sample target position in the test. For example, step 205 may include:
[0212] (205.1) Adjusting model parameters of an initial reinforcement learning model based on the target optimization reward to obtain a motion state of the sample first fiber locator; wherein the motion state includes a first state in which the sample first fiber locator successfully locates to the sample target position, a second state in which the sample first fiber locator collides with at least one adjacent sample second fiber locator after movement, and a third state in which the sample first fiber locator fails to locate and does not collide with an adjacent sample second fiber locator after movement;
[0213] (205.2) Obtaining a total number of samples of the plurality of sample first fiber locators participating in the reinforcement learning training, and when a first ratio of a first number of the plurality of sample first fiber locators in a first state, representing a motion state, to the total number of samples is lower than a first threshold, repeatedly adjusting model parameters of the initial reinforcement learning model based on the target optimization reward to obtain a new motion state of the sample first fiber locators;
[0214] (205.3) Until a first ratio of a first number of first fiber locators in a first state of a plurality of samples representing a motion state to a first total number of samples is higher than a first threshold, the adjustment of the model parameters of the initial reinforcement learning model is stopped to obtain a target reinforcement learning model.
[0215] The motion state may be a result state of the sample first optical fiber positioner after attempting to position the sample to the target posture.
[0216] The first state may be a state in which the first optical fiber locator of the sample is successfully positioned to the target posture of the sample.
[0217] The second state may be a state in which the sample first optical fiber positioner collides with at least one adjacent sample second optical fiber positioner after movement.
[0218] The third state may be a state in which the sample first optical fiber positioner is not successfully positioned after movement and does not collide with the adjacent sample second optical fiber positioner.
[0219] The total sample amount may be the total number of multiple sample first optical fiber locators participating in the reinforcement learning training.
[0220] The first number may be the number of sample first fiber locators in the first state after a training round.
[0221] The first ratio may be a ratio of a first number of first fiber locators in the first state to the total number of samples representing the motion state. First ratio = number of fiber locators in the first state / total number of samples.
[0222] The first threshold may be a preset ratio standard for determining whether to stop adjusting the model parameters. For example, the first threshold may be 80%, 85%, etc., and may be specifically set according to actual conditions.
[0223] Specifically, when the first ratio is higher than a first threshold, it indicates that the model has achieved accurate positioning, and therefore, training of the initial reinforcement learning model can be stopped. Furthermore, training of the initial reinforcement learning model can be stopped if the first ratio is higher than the first threshold for a consecutive number of times. The number of consecutive times can be set to 3, 5, or so.
[0224] In some embodiments, a quantity threshold may be set, for example, 50, 100, or the like. When the first quantity does not exceed the quantity threshold, reinforcement learning training of the initial reinforcement learning model continues; when the first quantity exceeds the quantity threshold, reinforcement learning training may be stopped to obtain a target reinforcement learning model. It is understood that the quantity threshold may be set based on actual circumstances.
[0225] For example, in the first round of training for the initial reinforcement learning model, assuming there are 10 sample first fiber positioners, 5 of which successfully reach the target rotation angle (first state), 2 collide with adjacent sample second manipulators or other objects during movement (second state), and 3 fail to locate successfully and do not collide (third state). In this case, the first ratio can be calculated as 5 / 10 = 0.5.
[0226] For example, if the set first threshold is 0.7, then the current ratio of 0.5 is lower than the first threshold, and the model parameters need to be further adjusted.
[0227] Furthermore, if the first ratio is lower than the first threshold during the second to fourth rounds, the model parameters are adjusted after each round and the initial reinforcement learning model is trained again. If, after the fifth round, the first fiber positioners of all nine samples successfully reach the target rotation angle, the first ratio is 0.9, which is higher than the first threshold. At this point, reinforcement learning training can be stopped, and the target reinforcement learning model can be obtained.
[0228] Through multiple rounds of operation and parameter adjustment, the model parameters are continuously optimized, so that the initial reinforcement learning model can learn to avoid actions that lead to collisions, thereby reducing collisions between fiber locators in the system and improving the stability and reliability of the system.
[0229] In some embodiments, to ensure that the trained target reinforcement learning model achieves optimal performance in practical applications, hyperparameter optimization can be performed on the model parameters to greatly improve the efficiency and effectiveness of model training and the accuracy and reliability of fiber positioning. For example, when adjusting the model parameters, the following adjustment methods can be used, including:
[0230] (A.1) Obtaining motion states corresponding to multiple training rounds; wherein each training round corresponds to an adjustment of model parameters;
[0231] (A.2) for each training round, obtaining a second ratio of a second number of first fiber positioners in the first state to the total number of samples representing the motion state;
[0232] (A.3) obtaining a second threshold value, and comparing the second ratio with the second threshold value to obtain a comparison result;
[0233] (A.4) assigning the first model parameters of the current training round to the corresponding probability model based on the comparison result, wherein the probability model includes the first probability distribution model and the second probability distribution model;
[0234] (A.5) calculating an expected improvement value in a corresponding probability model based on the first model parameters, and determining a second model parameter from the corresponding probability model based on the expected improvement value;
[0235] (A.6) The second model parameters are used as the updated model parameters of the initial reinforcement learning model.
[0236] The second ratio may refer to the ratio of the second number of sample first optical fiber positioners in the first state to the total number of samples in a training round, that is, the second ratio = the second number / the total number of samples.
[0237] The second threshold may be a preset standard for judging whether the training effect of the current training round is good or bad.
[0238] The comparison result may be a result of comparing the second ratio with the second threshold, for example, the second ratio is greater than the second threshold, the second ratio is less than the second threshold, etc.
[0239] The probability model may be a mathematical model used to describe and predict the relationship between model parameters and training effects.
[0240] The first probability distribution model may be a probability distribution for describing those hyperparameter configurations that result in successful training (ie, the second ratio is greater than the second threshold).
[0241] The second probability distribution model may be a probability distribution for describing those hyperparameter configurations that result in unsuccessful training (ie, the second ratio is less than the second threshold).
[0242] The second model parameter may be a model parameter used in a new round of training determined based on an expected improvement value calculated based on a probability model during hyperparameter optimization.
[0243] For example, if an initial reinforcement learning model is trained using a simulation with a focal plane containing 60 fiber positioners, each training round corresponds to an adjustment of the model parameters. For example, there are five training rounds of data: Round 1: motion state S1, Round 2: motion state S2, Round 3: motion state S3, Round 4: motion state S4, Round 5: motion state S5.
[0244] Furthermore, for each training round, a second ratio of the second number of the first fiber locator in the first state to the total number of samples representing the motion state is obtained. If the total number of samples in round 1 is 100 and the number of samples in the first state is 30, the second ratio is 0.3; for round 2, the total number of samples is 100 and the number of samples in the first state is 50, the second ratio is 0.5; for round 3, the total number of samples is 100 and the number of samples in the first state is 70, the second ratio is 0.7; for round 4, the total number of samples is 100 and the number of samples in the first state is 20, the second ratio is 0.2; and for round 5, the total number of samples is 100 and the number of samples in the first state is 40, the second ratio is 0.4.
[0245] Furthermore, if the second threshold is 0.4, the second ratio is compared with the second threshold, and the following results are obtained: the comparison result of round 1 is false; the comparison result of round 2 is true; the comparison result of round 3 is true; the comparison result of round 4 is false; and the comparison result of round 5 is true.
[0246] Furthermore, based on the comparison result, the first model parameters of the current training round are assigned to the corresponding probability model, wherein the first probability distribution model (the comparison result is true) includes the model parameters of rounds 2, 3 and 5, and the second probability distribution model (the comparison result is false) includes the model parameters of rounds 1 and 4.
[0247] Afterwards, the expected improvement value γ is calculated by the following formula * :
[0248]
[0249] in, Γis the set of all possible hyperparameter configurations, l(γ) is the first probability distribution model, g(γ) is the second probability distribution model, and γ is the hyperparameter configuration.
[0250] Furthermore, we can select a hyperparameter configuration that maximizes g(γ). During each round of training, we repeatedly determine the corresponding probability distribution model based on the comparison results and update the corresponding probability distribution model. This process is repeated until the optimal hyperparameter configuration, i.e., the second model parameters, is found. These second model parameters are then used as the optimal model parameters for the initial reinforcement learning model.
[0251] Through multiple training sessions and continuous adjustment of hyperparameter configurations, the initial reinforcement learning model can adapt to different training data and environments, improving its generalization capabilities. Furthermore, it can enable the initial reinforcement learning model to more accurately capture effective hyperparameter configurations, enabling efficient positioning of the fiber locator in complex environments and significantly improving the performance of the target reinforcement learning model in practical applications.
[0252] See also Figure 8 The present application also provides a control device for an optical fiber locator, which can implement the control method of the optical fiber locator. The control device for the optical fiber locator includes:
[0253] A first determining module 81 is configured to determine a plurality of first optical fiber positioners to be positioned, wherein each first optical fiber positioner has at least one rotatable first mechanical arm;
[0254] A second determining module 82 is configured to determine, for each first optical fiber positioner, at least one second optical fiber positioner adjacent to the first optical fiber positioner and a second mechanical arm corresponding to each second optical fiber positioner;
[0255] A first acquisition module 83 is used to acquire the current first pose of the first robotic arm and the target pose of the first robotic arm;
[0256] A second acquisition module 84 is configured to obtain the current second posture of each second robotic arm, and calculate a rotation path based on the first posture, target posture, and second posture of each adjacent second robotic arm using a target reinforcement learning model to obtain a target rotation path for each first robotic arm to avoid the adjacent second robotic arm;
[0257] A positioning module 85 is used to position the first robotic arm corresponding to the first fiber optic locator based on the target rotation route; wherein, the target reinforcement learning model is obtained after the target optimization reward adjustment model parameters are constructed by the initial reinforcement learning model based on the sample first pose, the sample first intermediate pose and the sample target pose; wherein, the sample first pose is the sample first pose of the sample first robotic arm of the sample first fiber optic locator, the sample first intermediate pose is the pose obtained by adjusting the sample first robotic arm according to the sample target rotation route corresponding to the perceived sample environment information, and the sample target pose is the target pose for the sample first robotic arm; the sample environment information at least includes the pose change information of the sample first intermediate pose of the sample first robotic arm of the sample first fiber optic locator, and the pose change information of the sample second pose of each sample second robotic arm corresponding to at least one sample second fiber optic locator adjacent to the sample first fiber optic locator.
[0258] The specific implementation of the control device of the optical fiber locator is basically the same as the specific embodiment of the control method of the optical fiber locator described above, and will not be repeated here. On the premise of meeting the requirements of the embodiment of the present application, the control device of the optical fiber locator can also be provided with other functional modules to implement the control method of the optical fiber locator in the above embodiment.
[0259] The present application also provides a computer device comprising a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the control method of the optical fiber locator. The computer device can be any intelligent terminal, including a tablet computer and an in-vehicle computer.
[0260] See also Figure 9 , Figure 9 The hardware structure of a computer device according to another embodiment is shown. The computer device includes:
[0261] The processor 91 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0262] The memory 92 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 92 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program codes are stored in the memory 92 and are called by the processor 91 to execute the control method of the optical fiber locator of the embodiments of this application.
[0263] Input / output interface 93, used for information input and output;
[0264] Communication interface 94, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, Wi-Fi, Bluetooth, etc.);
[0265] bus 95 , which transmits information between the various components of the device (e.g., processor 91 , memory 92 , input / output interface 93 , and communication interface 94 );
[0266] The processor 91 , the memory 92 , the input / output interface 93 and the communication interface 94 are connected to each other in communication within the device via a bus 95 .
[0267] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, which implements the control method of the optical fiber locator when executed by a processor.
[0268] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0269] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0270] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0271] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0272] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0273] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0274] It should be understood that in this application, "at least one (item)" and "several" refer to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0275] In the several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the above units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0276] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0277] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0278] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store programs.
[0279] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A control method for an optical fiber positioner, characterized in that: The method comprises: determining a plurality of first optical fiber positioners to be positioned, wherein each of the first optical fiber positioners has at least one rotatable first mechanical arm; For each of the first optical fiber positioners, determining at least one second optical fiber positioner adjacent to the first optical fiber positioner and a second robotic arm corresponding to each second optical fiber positioner; Obtaining a current first pose of the first robotic arm and a target pose for the first robotic arm; Obtain the current second pose of each second robotic arm, and calculate the rotation path based on the first pose of each first robotic arm, the target pose, and the second pose of each adjacent second robotic arm through the target reinforcement learning model to obtain the target rotation path of each first robotic arm to avoid the adjacent second robotic arm; Positioning the first robotic arm corresponding to the first optical fiber positioner based on the target rotation route; The target reinforcement learning model is obtained by adjusting the model parameters of the target optimization reward constructed by the initial reinforcement learning model based on the sample first pose, the sample first intermediate pose and the sample target pose; Among them, the sample first posture is the sample first posture of the sample first robotic arm of the sample first optical fiber locator, the sample first intermediate posture is the posture obtained by adjusting the sample first robotic arm according to the sample target rotation route corresponding to the perceived sample environment information, and the sample target posture is the target posture for the sample first robotic arm; the sample environment information at least includes the posture change information of the sample first intermediate posture of the sample first robotic arm of the sample first optical fiber locator, and the posture change information of the sample second posture of each sample second robotic arm corresponding to at least one sample second optical fiber locator adjacent to the sample first optical fiber locator.
2. The control method of the optical fiber positioner according to claim 1, characterized in that: The target reinforcement learning model is trained in the following way: obtaining a sample first position of at least one sample first mechanical arm of a sample first optical fiber positioner; Perceiving sample environment information corresponding to the first optical fiber locator of the sample through an initial reinforcement learning model, and obtaining a sample target rotation path; Adjusting the posture of each sample first manipulator according to the sample target rotation route, and determining the sample first intermediate posture after the adjustment of the sample first manipulator; Obtaining a sample target pose for each of the sample first robotic arms, and constructing a target optimization reward based on the sample first pose, the sample first intermediate pose, and the sample target pose; The model parameters of the initial reinforcement learning model are adjusted according to the target optimization reward to obtain a target reinforcement learning model.
3. The control method of the optical fiber positioner according to claim 2, characterized in that: The constructing of the target optimization reward based on the sample first pose, the sample first intermediate pose, and the sample target pose includes: At a first sample time, determining a collision reward based on the distance between the sample first fiber locator and a plurality of adjacent sample second fiber locators; wherein the first sample time is the time after the initial reinforcement learning model adjusts the sample first manipulator's sample target rotation path based on the sample environment information perceived; determining a target distance reward based on a distance between at least one sample first manipulator of each sample first optical fiber positioner and a sample target posture after the at least one sample first manipulator of each sample first optical fiber positioner rotates from the sample first posture to the sample first intermediate posture; Determining a time reward based on the interaction time between the current initial reinforcement learning model and the sample environment information; Determining a completion indicator reward based on the positioning status of the plurality of sample first optical fiber locators; Determine a target optimization reward of the initial reinforcement learning model at the first sample time based on the collision reward, the target distance reward, the time reward, and the completion index reward.
4. The control method of the optical fiber positioner according to claim 3, characterized in that: The determining of the collision reward based on the distance between the sample first optical fiber locator and a plurality of adjacent sample second optical fiber locators at the first sample time comprises: At a first sample time, for each of the sample first optical fiber locators, obtaining a minimum sample distance between the profiles of the sample first optical fiber locator and each adjacent sample second optical fiber locator; Accumulating multiple minimum sample distances between each of the sample first optical fiber locators and multiple adjacent sample second optical fiber locators to determine a first sample distance; For a plurality of sample first manipulators of a first optical fiber positioner of a plurality of samples to be positioned, obtaining a plurality of corresponding first sample distances; Based on the multiple first sample distances, a collision reward of the initial reinforcement learning model at the first sample time is determined.
5. The control method of the optical fiber positioner according to claim 3, characterized in that: The target distance reward is determined based on the distance between at least one sample first manipulator of each sample first optical fiber positioner and the sample target posture after the at least one sample first manipulator rotates from the sample first posture to the sample first intermediate posture, including: At a first sample moment, determining a sample first intermediate posture obtained by adjusting each of the sample first mechanical arms of the sample first optical fiber positioner from a sample first posture according to a sample target rotation route corresponding to the sensed sample environment information; At the second sample moment, determining a sample second intermediate pose obtained by adjusting each of the sample first manipulators from the sample first intermediate pose according to a sample target rotation route corresponding to the perceived sample environment information; Acquire a sample target posture that each of the sample first mechanical arms of the sample first optical fiber positioner needs to position; For each of the sample first mechanical arms of each of the sample first optical fiber positioners, obtaining a first difference in rotation angle between the sample first posture and the sample target posture, and obtaining a second difference in rotation angle between the sample second posture and the sample target posture; determining a first target distance reward for each of the sample first fiber locators based on the first difference and the second difference of each of the sample first robotic arms; For a plurality of sample first optical fiber locators to be located, a plurality of corresponding first target distance rewards are obtained, and based on the plurality of first target distance rewards, a target distance reward of the initial reinforcement learning model is determined.
6. The control method of the optical fiber positioner according to claim 2, characterized in that: The adjusting the model parameters of the initial reinforcement learning model according to the target optimization reward to obtain a target reinforcement learning model includes: The model parameters of the initial reinforcement learning model are adjusted based on the target optimization reward to obtain the motion state of the sample first optical fiber locator; wherein the motion state includes a first state when the sample first optical fiber locator is successfully positioned to the sample target posture, a second state when the sample first optical fiber locator collides with at least one adjacent sample second optical fiber locator after movement, and a third state when the sample first optical fiber locator is not successfully positioned after movement and does not collide with the adjacent sample second optical fiber locator; Obtaining a total number of samples of a plurality of sample first optical fiber locators participating in reinforcement learning training, and when a first ratio of a first number of the plurality of sample first optical fiber locators in a first state, characterized by the motion state, to the total number of samples is lower than a first threshold, repeatedly adjusting model parameters of the initial reinforcement learning model according to the target optimization reward to obtain a new motion state of the sample first optical fiber locator; Until the first ratio of the first number of the first fiber locators in the first state of the motion state characterizing multiple samples to the total number of samples is higher than the first threshold, the adjustment of the model parameters of the initial reinforcement learning model is stopped to obtain the target reinforcement learning model.
7. The control method of the optical fiber positioner according to claim 6, characterized in that: The method further comprises: Obtaining motion states corresponding to multiple training rounds, where each training round corresponds to an adjustment of model parameters; For each of the training rounds, obtaining a second ratio of a second number of the first optical fiber positioners in the first state of the plurality of samples representing the motion state to the total number of the samples; Obtaining a second threshold, and comparing the second ratio with the second threshold to obtain a comparison result; Allocating the first model parameters of the current training round to the corresponding probability model based on the comparison result, wherein the probability model includes a first probability distribution model and a second probability distribution model; calculating an expected improvement value in a corresponding probability model based on the first model parameter, and determining a second model parameter from the corresponding probability model based on the expected improvement value; The second model parameters are used as updated model parameters of the initial reinforcement learning model.
8. A control device for an optical fiber positioner, characterized in that: The device comprises: a first determining module, configured to determine a plurality of first optical fiber positioners to be positioned, wherein each of the first optical fiber positioners has at least one rotatable first mechanical arm; a second determining module configured to determine, for each of the first optical fiber positioners, at least one second optical fiber positioner adjacent to the first optical fiber positioner and a second mechanical arm corresponding to each second optical fiber positioner; A first acquisition module is used to acquire a current first pose of the first robotic arm and a target pose for the first robotic arm; A second acquisition module is used to obtain the current second posture of each second robotic arm, and calculate the rotation path based on the first posture, target posture and the second posture of each adjacent second robotic arm of each first robotic arm through the target reinforcement learning model to obtain the target rotation path of each first robotic arm in avoiding the adjacent second robotic arm; A positioning module for positioning the first robotic arm corresponding to the first optical fiber locator based on the target rotation route; wherein the target reinforcement learning model is obtained after the target optimization reward adjustment model parameters are constructed by the initial reinforcement learning model based on the sample first pose, the sample first intermediate pose and the sample target pose; wherein the sample first pose is the sample first pose of the sample first robotic arm of the sample first optical fiber locator, the sample first intermediate pose is the pose obtained by adjusting the sample first robotic arm according to the sample target rotation route corresponding to the perceived sample environment information, and the sample target pose is the target pose for the sample first robotic arm; the sample environment information at least includes the pose change information of the sample first intermediate pose of the sample first robotic arm of the sample first optical fiber locator, and the pose change information of the sample second pose of each sample second robotic arm corresponding to at least one sample second optical fiber locator adjacent to the sample first optical fiber locator.
9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the control method of the optical fiber locator according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the control method of the optical fiber positioner according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
High-voltage cable target identification and positioning method and device
CN113807348A
Mobile phone flexible flat cable positioning method based on visual touch information fusion
CN116630417A