Pipeline design system
The pipeline design system uses reinforcement learning to optimize pipe types and quantities, addressing inefficiencies in existing systems by designing high-quality underground pipeline routes that mimic experienced designer input and enable efficient construction.
Patent Information
- Application Number
- JP2023214919
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-20
- Publication Date
- 2025-07-02
AI Technical Summary
Existing pipeline design systems fail to consider the optimal combination of pipe types and quantities when automatically designing underground pipelines, leading to inefficiencies and potential design mistakes due to reliance on designer experience.
A pipeline design system utilizing reinforcement learning to determine the optimal route and combination of pipe types and quantities by setting a three-dimensional underground design space and employing reinforcement learning to maximize preset rewards, considering factors like separation distance from buried objects and arrival at the endpoint.
The system efficiently designs high-quality pipeline routes by optimizing pipe types, quantities, and combinations, mimicking experienced designer input and enabling quantitative comparison of design options, facilitating easier construction and information sharing.
Smart Images

Figure 2025098644000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a pipeline design system for designing pipelines that can accommodate wires and the like inside underground.
Background Art
[0002] A large number of pipelines formed by sequentially connecting a plurality of pipes are buried underground, and wires and the like are disposed in each of these pipelines. Conventionally, for each pipeline, it has generally been carried out manually by a designer using an information processing terminal or the like, while bypassing the areas of existing buried objects between the designated starting point and ending point on the drawing. There have been problems such as the quality of the designed pipeline depending on the designer's experience value, and a lot of time being spent on the work due to rework caused by design mistakes.
[0003] In order to address these issues, in recent years, systems for automatically designing each pipeline have been considered. For example, the pipeline design program disclosed in Patent Document 1 calculates a line extending in a trapezoidal shape that bypasses the area where existing buried objects are present on the longitudinal section data between the designated starting point and ending point, and then sequentially obtains the coordinate values of the intersection points of the line extending from the starting point to the ending point and the generated trapezoidal line. After that, it outputs the broken line obtained by sequentially connecting the obtained coordinate values from the starting point to the ending point as the optimal route of the pipeline.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] By the way, each pipe that makes up the pipeline includes both straight pipes and curved pipes, and furthermore, there are multiple types of pipes with different lengths and curvatures respectively. Therefore, when designing the pipeline, it is necessary to consider how to connect these multiple types of pipes. In addition to the types and quantities of pipes required for the construction, it is necessary to grasp their combinations in advance. However, in Patent Document 1, although the optimal route of the pipeline is automatically output, no consideration is given to the design of what types and quantities of pipe combinations the pipeline in the output route is composed of.
[0006] The present invention has been made in view of such points, and the object thereof is to provide a pipeline design system that can optimally and efficiently automatically design not only the route of the pipeline but also multiple types of pipes constituting the pipeline, their respective quantities, and their combinations.
Means for Solving the Problems
[0007] In order to achieve the above object, the present invention is characterized in that a combination of multiple types of pipes constituting a pipeline in a predetermined section can be acquired by reinforcement learning. Specifically, targeting a pipeline design system capable of designing a pipeline configured by sequentially connecting multiple pipes underground and extending around one or more existing buried objects, the following measures were taken.
[0008] That is, the pipeline design system according to the first invention includes a design space setting unit that sets a three-dimensional underground design space, and selects the type of the pipe and its extending direction as actions, and uses the data of each pipe selected by the actions as a state where the leading ends are sequentially connected, and a reinforcement learning unit that learns and outputs one or more pieces of learned pipeline data with the largest total of preset rewards obtained based on a series of the actions when it is possible to move from a specified starting point to an ending point in the underground design space. In the pipeline design system configured as described above, while bypassing existing buried objects from a specified starting point to an ending point underground, an optimal route of the pipeline extending is designed, and at the same time, it acts so as to find the optimal type, the respective number, and their combination for a plurality of pipes constituting the pipeline in the optimal route.
[0009] In the pipeline design system of the second invention, in the first invention, the reward is a preset score according to at least one of the separation distance from the data of the existing buried object, the relative distance to the ending point, the arrival at the ending point, and the number of actions. In the pipeline design system configured as described above, the learning pipeline data output by reinforcement learning acts so as to approach that by the design performed by a designer with high experience value by the conventional method.
[0010] The pipeline design system of the third invention is characterized in that, in the first or second invention, it includes a feature extraction unit that respectively calculates and extracts feature values of a plurality of the learning pipeline data, and an evaluation value output unit that converts each of the feature values into an evaluation value based on a preset evaluation table and outputs them respectively. In the pipeline design system configured as described above, the features of a plurality of learning pipeline data obtained by reinforcement learning can be understood, and it acts so that a designer can quantitatively compare each learning pipeline data.
[0011] The pipeline design system of the fourth invention is characterized in that, in the third invention, the feature value is at least one of the total length of the learning pipeline data, the average buried depth from the ground surface of the learning pipeline data, and the number of times the learning pipeline data passes under the existing buried object. In the pipeline design system configured as described above, it acts so that a plurality of learning pipeline data obtained by reinforcement learning can be quantitatively compared and considered in accordance with points important in pipeline design.
[0012] In the pipeline design system of the fifth invention, in the first or second invention, a drawing output unit is provided that outputs, respectively, a planar pipeline data diagram of the underground design space in which the learning pipeline data is drawn as viewed from above, and a longitudinal section pipeline data diagram as viewed from a predetermined side. In the pipeline design system configured in this way, it functions so that the route of the created pipeline can be displayed on two drawings.
Advantages of the Invention
[0013] In the pipeline design system of the first invention, while designing the optimal route of a pipeline that extends from a starting point specified underground to an end point while bypassing existing buried objects, the optimal type, the respective number, and their combinations of a plurality of pipes constituting the pipeline in the optimal route can be known. Therefore, the design of pipelines underground can be efficiently performed.
[0014] In the pipeline design system of the second invention, since the reward used for reinforcement learning is set in accordance with points that are important in pipeline design, the learning pipeline data output by reinforcement learning, for example, approaches that by a designer with a high level of experience using conventional methods. Therefore, it is possible to efficiently output a high-quality pipeline route, the types and respective numbers of pipes constituting the pipeline, and their combinations.
[0015] In the pipeline design system of the third invention, since the characteristics of a plurality of learning pipeline data obtained by reinforcement learning can be known, designers can quantitatively compare each learning pipeline data. Therefore, designers can find the highest-quality pipeline among a plurality of learning pipeline data.
[0016] In the pipeline design system of the fourth invention, a plurality of learned pipeline data obtained by reinforcement learning can be quantitatively compared and considered in accordance with the points important in pipeline design. Therefore, it is possible to obtain the route of the highest-quality pipeline, and it is possible to find high-quality types and numbers of pipes and their combinations in that route.
[0017] In the pipeline design system of the fifth invention, the route of the created pipeline can be displayed on two drawings. Therefore, for example, it becomes easier to give instructions when constructing the pipeline on-site, or it becomes easier to share pipeline information with others.
Brief Description of the Drawings
[0018]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Embodiments for Carrying Out the Invention
[0019] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. It should be noted that the following description of the preferred embodiments is merely illustrative in nature.
[0020] FIG. 1 shows a block diagram of a pipeline design system 1 according to an embodiment of the present invention. As shown in FIG. 2, this pipeline design system 1 automatically designs a pipeline 2 in which wires and the like are disposed inside the ground, and includes a computer 10 capable of outputting an optimal route in the ground, a combination of each pipe 3 constituting the pipeline 2, etc. by reinforcement learning and evaluation determination processing.
[0021] Reinforcement learning is, as shown in Figure 3, a method of learning a strategy for enabling an agent to take an optimal action a according to the situation at that time in an environment. In the present invention, the reinforcement learning unit 4b shown in Figure 1 corresponds to the agent, and by repeatedly executing the action a in a given environment and receiving the state s' and the reward r from the transitioned environment, the matrix (Q-table) for determining the action a is updated, and finally, it is a method of learning the action a for maximizing the reward r.
[0022] As shown in Figure 1, the computer 10 includes an information processing unit 4 that performs reinforcement learning and evaluation determination processing, an input database 5 that stores data input to the information processing unit 4 using a keyboard, a mouse, etc., a pipeline design database 6 that stores various data used for reinforcement learning and evaluation determination processing, and an output database 7 that stores each processing result obtained by being processed by the information processing unit 4.
[0023] As shown in Figures 4 to 6, the input database 5 stores a subsurface design space 5A that structures the location where the pipeline 2 to be designed is buried into structured data, design route condition data 5B that sets the conditions necessary for the design of the pipeline 2, and learning condition data 5C that sets the number of route search times (number of times of action a) i and the number of learning times k for reinforcement learning performed during the design of the pipeline 2. As shown in Figures 4 and 5, the subsurface design space 5A is defined by dividing the location where the pipeline 2 is buried into a large number of block-shaped small regions in an orthogonal coordinate system defined by the X-axis, Y-axis, and Z-axis. For example, each small region corresponding to the position where there are buried objects such as culverts and waterways is made to have information indicating the existence of these buried objects. The subsurface design space 5A has a plurality of existing buried object data 5a representing culverts, waterways, etc., position data 5b of the start point S and end point G of the pipeline 2 to be designed, and position data 5c of the ground surface F. In Figure 5, for the sake of convenience, the area of the existing buried object data 5a is shown by hatching.
[0024] The design route condition data 5B is setting parameters that are input and saved during the design of pipeline 2, as shown in Fig. 6. Numerical values are saved for the pipeline diameter of pipe 3 used during the construction of pipeline 2, the minimum separation distance from each existing buried object of the designed pipeline 2, and the required burial depth of pipeline 2, each adjusted according to the construction conditions. For example, a numerical value is saved for selecting one from six options for the pipeline diameter of each pipe 3 that makes up the designed pipeline 2, a numerical value is saved for selecting one from two options for the minimum separation distance from each existing buried object of the designed pipeline 2, and a numerical value is saved for selecting one from three options for the required burial depth of the designed pipeline 2. And as shown in Fig. 5, the area 5d of the minimum separation distance and the area 5e of the required burial depth are saved as data in the subsurface design space 5A.
[0025] The learning condition data 5C is setting parameters that are input and saved during the design of pipeline 2. Numerical values are saved for the number of route search times (number of actions a) i and the number of learning times k for the reinforcement learning performed during the design of pipeline 2, each adjusted according to the design conditions. For example, when it is desired to increase the number of route search times i for the designed pipeline 2 in the reinforcement learning performed by the pipeline design system 1, the arbitrarily input number with the upper limit value set to 100,000 times is saved, and when it is desired to decrease it, the arbitrarily input number with the lower limit value set to 100 times is saved. Also, for the pipeline 2 designed in the reinforcement learning performed by the pipeline design system 1, when a large number of route patterns are to be considered, the arbitrarily input number with the upper limit value set to 100,000 times is saved, and when it is desired to decrease the number, the arbitrarily input number with the lower limit value set to 100 times is saved.
[0026] As shown in Fig. 1, the pipeline design database 6 stores the reinforcement learning data 6A and the evaluation table 6B. The reinforcement learning data 6A is composed of the action data 6a and the reward data 6b necessary for performing reinforcement learning. As shown in Fig. 7(a), the action data 6a stores the selection pattern of action a in reinforcement learning. In an embodiment of the present invention, when selecting a straight pipe that extends linearly for the pipe 3 selected during action a, there is one option. When selecting a curved pipe that extends with a bend, there are 28 options as there are 7 types with different curvatures and 4 extending directions for each type. In total, 29 selection patterns are stored.
[0027] As shown in Fig. 7(b), the reward data 6b stores the points given as a pre-determined reward r according to the environment and situation after the transition in reinforcement learning. For example, -1 point is given as the reward r every time action a is performed. Also, when the relative distance to the end point G decreases after performing action a, +2 points are given as the reward r. Further, when the distance from the existing buried object falls below a pre-determined minimum separation distance after performing action a, -10 points are given as the reward r. And when arriving at the end point G after performing action a, +100 is given as the reward r.
[0028] In the evaluation table 6B, as shown in FIG. 8, the evaluation items, evaluation methods, and ways of comprehensive evaluation when evaluating each learning pipeline data 20 obtained by reinforcement learning are stored. For example, the scores assigned according to the total length of the learning pipeline data 20 are stored. Also, the method of assigning scores by comparing the average excavation depths of each learning pipeline data 20 with each other is stored. Furthermore, the method of assigning scores by comparing the number of protective measures of each learning pipeline data 20 with each other is stored. For example, the method of calculating the total length of the learning pipeline data 20 and evaluating it with 1 point for every 1 m of its total length is stored. Also, the method of calculating the average excavation depth of each learning pipeline data 20 and evaluating it as 1 point, 2 points,..., n points (n is a natural number) in the order of the learning pipeline data 20 with shallower depths is stored. Furthermore, the method of evaluating each learning pipeline data 20 as 1 point, 2 points,..., n points (n is a natural number) in the order of those with fewer protective measures is stored. Then, the scores evaluated for each evaluation item are totaled for each learning pipeline data 20 as an evaluation value, and the evaluation method of selecting the learning pipeline data 20 with the lowest evaluation value as the one that can construct the best pipeline 2 is stored.
[0029] As shown in FIG. 1, the information processing unit 4 is composed of a design space setting unit 4a, a reinforcement learning unit 4b, a feature extraction unit 4c, an evaluation value output unit 4d, a drawing output unit 4e, a material quantity output unit 4f, a tension calculation unit 4g, and a lateral pressure calculation unit 4h. As shown in FIGS. 4 and 5, the design space setting unit 4a is configured to set the underground design space 5A in the computer 10.
[0030] The reinforcement learning unit 4b takes the selection of the type of pipe 3 and its extending direction as the action a, and takes the state s' after transitioning to the end point connected in sequence with the data of each pipe 3 selected by the action a as the state s'. Based on a series of each action a when it can move from the designated starting point S to the ending point G in the underground design space 5A, it learns and outputs one or more learning pipeline data 20 with the most total preset rewards r obtained respectively by reinforcement learning.
[0031] Specifically, for example, when the reinforcement learning unit 4b selects "a 1m straight pipe" (hereinafter referred to as this pipe 3 as the first selected pipe data 3a) from the selection pattern of the action data 6a at the starting point S, since the extending direction is only "straight line", the first selected pipe data 3a is arranged to extend straight toward the end point G, and the tip position of the extended first selected pipe data 3a is set as the transition state s'. At this time, a reward r can be obtained based on the reward data 6b at the extended position of the first selected pipe data 3a. While 1 point is subtracted because the action a has been performed once, 2 points are added because the relative distance to the end point G has become shorter, and the reward r is increased by 1 point in total. Next, the reinforcement learning unit 4b randomly selects the type of the pipe 3 from the selection pattern of the action data 6a. If "pipe type = curved pipe (R5)", "extending direction = upward" (hereinafter referred to as this pipe 3 as the second selected pipe data 3b) is selected as the type of the pipe 3, the second selected pipe data 3b is connected to the first selected pipe data 3a so as to extend and bend upward, and the tip position of the extended second selected pipe data 3b is set as the transition state s'. At this time, a reward r can be obtained based on the reward data 6b at the extended position of the second selected pipe data 3b. While 1 point is subtracted because the action a has been performed once, 2 points are added because the relative distance to the end point G has become shorter, and the reward r is increased by 1 point in total.
[0032] Next, the reinforcement learning unit 4b randomly selects the type of pipe 3 from the selection patterns of the action data 6a. If, as the type of pipe 3, "pipe type = curved pipe (R15)" and "extending direction = downward" (hereinafter, this pipe 3 is referred to as the third selection pipe data 3c) are selected, the third selection pipe data 3c is connected to the second selection pipe data 3b so as to bend downward and extend, and the state s' is set to the state where the tip position of the extended third selection pipe data 3c is transitioned. At this time, a reward r is obtained based on the reward data 6b at the extended position of the third selection pipe data 3c. While one point is subtracted because the action a has been performed once, two points are added because the relative distance to the end point G has become shorter, and the reward r is increased by one point in total. Here, for example, if the extended position of the third selection pipe data 3c falls below the minimum separation distance from the existing buried object data 5a, the reward r is decreased by 10 points. That is, when the route of the pipeline 2 to be designed does not bypass the existing buried object, the reward r is reduced. In this way, the reinforcement learning unit 4b sequentially connects the data of the randomly selected pipe 3 while selecting the action a and obtaining the reward r. For example, as shown in FIG. 11, one or more learning pipeline data 20 (learning pipeline data 20A, 20B, 20C) that reach the end point G are obtained, and the learning pipeline data 20 with a high total reward r has a route that bypasses the existing buried object and extends.
[0033] The feature extraction unit 4c calculates and extracts the feature values of each learning pipeline data 20 obtained by the reinforcement learning unit 4b. As shown in FIG. 8, the feature values are the total length of each learning pipeline data 20, the average burial depth from the ground surface F of each learning pipeline data 20, and the number of times each learning pipeline data 20 passes below the existing buried object, that is, the number of times it is necessary to take protective measures. For example, for the learning pipeline data 20A in FIG. 11, the total length is 50 m, and the average buried depth is 8 m. Since its route passes above the culvert and below the water pipe, the number of times protection measures need to be taken is extracted as 1. Also, for the learning pipeline data 20B in FIG. 11, the total length is 52 m, and the average buried depth is 6 m. Since its route passes above the culvert and the water pipe, it is extracted as not requiring protection measures. Further, for the learning pipeline data 20C in FIG. 11, the total length is 56 m, and the average buried depth is 8 m. Since its route passes below the culvert and also below the water pipe, the number of times protection measures need to be taken is extracted as 2.
[0034] The evaluation value output unit 4d converts each of the above-described characteristic values into an evaluation value based on a predetermined evaluation table 6B and outputs them respectively. In the embodiment of the present invention, for the evaluation value of the length of the learning pipeline data 20, 1 m is converted as 1 point. Also, the evaluation values are converted as 1 point, 2 points,..., n points (n is a natural number) in the order of the learning pipeline data 20 with a shallower average buried depth from the ground surface F. Further, the evaluation values are converted as 1 point, 2 points,..., n points (n is a natural number) in the order of the learning pipeline data 20 that requires less protection measures. Then, the evaluation values of each learning pipeline data 20 are summed up, and the learning pipeline data 20 with the lowest total score is selected as the learning pipeline data 20 with the highest evaluation. As shown in FIG. 12, for example, the overall evaluations of the learning pipeline data 20A, 20B, and 20C shown in FIG. 11 by the evaluation value output unit 4d are 54 points, 54 points, and 62 points respectively. The learning pipeline data 20A or the learning pipeline data 20B is selected as the learning pipeline data 20 with the highest evaluation and stored in the output database 7 as the route evaluation data 7g.
[0035] As shown in FIGS. 13 and 14, the drawing output unit 4e outputs a planar pipeline data diagram 7a of the underground design space 5A in which the learning pipeline data 20 selected by the evaluation value output unit 4d is drawn, as viewed from above, and a longitudinal section pipeline data diagram 7b as viewed from a predetermined side, and stores them in the output database 7. Also, as shown in FIG. 15, the drawing output unit 4e outputs connection material diagram data 7c in which the combinations, quantities, and length dimensions of the respective pipes 3 constituting the learning pipeline data 20 selected by the evaluation value output unit 4d are described, and stores them in the output database 7.
[0036] The material quantity output unit 4f calculates the types and respective quantities of the respective pipes 3 constituting the learning pipeline data 20 selected by the evaluation value output unit 4d as material quantity data 7d, and stores it in the output database 7. The tension calculation unit 4g performs a pulling tension calculation necessary for cable laying in the learning pipeline data 20 selected by the evaluation value output unit 4d. Specifically, the tension calculation unit 4g calculates the pulling tension for each change point of the learning pipeline data 20 based on, for example, a conventionally well-known calculation formula described in Japanese Patent Application Laid-Open No. 2013-54450, outputs the tension calculation data 7e, and stores it in the output database 7. The lateral pressure calculation unit 4h performs a lateral pressure calculation of the learning pipeline data 20 selected by the evaluation value output unit 4d. Specifically, the lateral pressure calculation unit 4h calculates the lateral pressure of the pipe 3 composed of the curved pipes of the learning pipeline data 20 based on, for example, a conventionally well-known calculation formula described in Japanese Patent Application Laid-Open No. 2013-54450, outputs the lateral pressure calculation data 7f, and stores it in the output database 7.
[0037] Next, the operation of outputting the learning pipeline data 20 with the highest evaluation in the information processing unit 4 of the pipeline design system 1 according to the embodiment of the present invention will be described in detail. As shown in FIG. 9, first, in step S1, after inputting the underground design space 5A in which the location where the pipeline 2 to be designed is buried is made into structured data to the computer 10, the process proceeds to step S2. In step S2, design route condition data 5B is input to the computer 10. Specifically, for the computer 10, numerical values are input according to the construction conditions for the pipe path of the pipe 3 used in the pipeline 2, the minimum separation distance from each existing buried object of the pipeline 2 to be designed, and the required burial depth of the pipeline 2, respectively.
[0038] After inputting the design route condition data 5B in step S2, the process proceeds to step S3, and learning condition data 5C is input to the computer 10. Specifically, for the computer 10, numerical values m and n (m and n are natural numbers) are input according to the design conditions for the number of route search times (action times) i and the number of learning times k performed during the design of the pipeline 2.
[0039] After inputting the learning condition data 5C in step S3, the process proceeds to step S4, where the learning number k is initialized and 1 is assigned to the number k, and then the process proceeds to step S5. Also, when proceeding to step S5, the state s in the reinforcement learning is initialized (the number of route search times is initialized), and 1 is assigned to the number i, and then the process proceeds to step S6.
[0040] In step S6, the reinforcement learning unit 4b randomly selects one from each selection pattern of the action data 6a. For example, if the first selection pipe data 3a of "1m straight pipe" is selected from the selection patterns of the action data 6a, the first selection pipe data 3a is arranged at the starting point S so as to extend straight toward the end point G, and the process proceeds to step S7. The tip position of the extended first selection pipe data 3a is set as the transition state s', and a reward r is received based on the reward data 6b. At this time, the received reward r is decreased by 1 point because the action a is performed once, while 2 points are added because the relative distance to the end point G becomes shorter, and the total reward r is increased by 1 point.
[0041] When receiving the reward r in step S7, the process proceeds to step S8 to determine whether the state s' has reached the end point G. When the determination in step S8 is YES, that is, when it is determined that the state s' has reached the end point G, the learning pipeline data 20 is obtained, and then the process proceeds to step S10.
[0042] On the other hand, when the determination in step S7 is NO, that is, when it is determined that the state s' has not reached the end point G, the process proceeds to step S9 to determine whether the number of actions i has reached m times. When the determination in step S9 is YES, that is, when the number of actions i has reached m times, the process returns to step S5 to start the action a from the beginning. On the other hand, when the determination in step S9 is NO, that is, when the number of actions i has not reached m times, the process proceeds to step S11, substitutes i + 1 for i, and then returns to step S6 to process the next action a. For example, when the reinforcement learning unit 4b selects the second selection pipe data 3b of "pipe type = curved pipe (R5)" and "extending direction = up" from each selection pattern of the action data 6a, the second selection pipe data 3b is connected to the first selection pipe data 3a so as to bend upward and extend, and the process proceeds to step S7. The reward r is obtained based on the reward data 6b with the tip position of the extended second selection pipe data 3b as the transition state s'.
[0043] When the process proceeds to step S10, the computer 10 determines whether the number of learning times k has reached n times. When the determination in this step S10 is NO, that is, when the number of learning times k is not n times, the process proceeds to step S12 to save the learning pipeline data 20 in the output database 7. Then, the process proceeds to step S13, substitutes k + 1 for k, and then returns to step S5 to start the action a from the beginning. On the other hand, when the determination in step S10 is YES, that is, when the number of learning times k reaches n times, as shown in FIG. 10, the process proceeds to step S14, and an evaluation value is output for each of the output learning pipeline data 20. First, the feature extraction unit 4c calculates the total length, the average burial depth from the ground surface F, and the number of times each learning pipeline data 20 passes below the existing buried object for each of the learning pipeline data 20 obtained by the reinforcement learning unit 4b. Specifically, assuming that three learning pipeline data 20A, 20B, and 20C in FIG. 11 are obtained by the reinforcement learning unit 4b, the feature extraction unit 4c, for the learning pipeline data 20A, for example, the total length is 50 m, and the average burial depth is 8 m, and its route passes above the culvert while passing below the water pipe, so the number of times of taking protective measures is extracted as 1. The feature values are extracted in the same way for the learning pipeline data 20B and 20C. After that, the evaluation value output unit 4d converts the evaluation values of the learning pipeline data 20A, 20B, and 20C based on the evaluation method shown in FIG. 8 for each of the above-obtained feature values, and outputs a comprehensive evaluation as shown in FIG. 12.
[0044] In step S14, when the evaluation of each learning pipeline data 20 is completed, the process proceeds to step S15, and the drawing output unit 4e outputs the plan pipeline data diagram 7a and the longitudinal section pipeline data diagram 7b in the learning pipeline data 20 with the highest evaluation selected by the evaluation value output unit 4d, and proceeds to step S16. In step S16, the drawing output unit 4e outputs the connection material diagram data 7c showing the combination, quantity, and length dimensions of each pipe 3 constituting the learning pipeline data 20 with the highest evaluation selected by the evaluation value output unit 4d, and proceeds to step S17.
[0045] In step S17, the material quantity output unit 4f calculates the types and quantities of each pipe 3 constituting the learning pipeline data 20 with the highest evaluation selected by the evaluation value output unit 4d, and proceeds to step S18. In step S18, the tension calculation unit 4g performs a draw-in tension calculation required for cable laying on the learning pipeline data 20 selected by the evaluation value output unit 4d, and outputs tension calculation data 7e. Further, the side pressure calculation unit 4h performs a side pressure calculation on the learning pipeline data 20 selected by the evaluation value output unit 4d, outputs side pressure calculation data 7f, and then ends the processing of the information processing unit 4.
[0046] As described above, according to the embodiment of the present invention, while bypassing existing buried objects such as culverts and water pipes from the specified starting point S to the ending point G in the ground, the optimal route of the pipeline 2 extending is designed, and at the same time, the optimal type, the respective number, and the combination thereof of the plurality of pipes 3 constituting the pipeline 2 in the optimal route can be known. Therefore, the design of the pipeline 2 in the ground can be efficiently performed. In addition, since the reward r used for reinforcement learning when designing the pipeline 2 of the optimal route is set in accordance with the points important in designing the pipeline 2, the learning pipeline data 20 output by reinforcement learning approaches, for example, the design by a designer with high experience values. Therefore, it is possible to efficiently output the route of the pipeline 2 with good quality, the types and the respective numbers of the pipes 3 constituting the pipeline 2, and the combination thereof.
[0047] In addition, since the characteristics of the plurality of learning pipeline data 20 obtained by reinforcement learning can be understood by the numerical values output by the feature extraction unit 4c, the designer can quantitatively compare each learning pipeline data 20. Therefore, the designer can find the pipeline 2 with the best quality among the plurality of learning pipeline data 20. In addition, since the feature values output by the feature extraction unit 4c are set in accordance with the points important in designing the pipeline 2, the plurality of learning pipeline data 20 obtained by reinforcement learning can be quantitatively compared and studied in accordance with the points important in designing the pipeline 2. Therefore, it is possible to obtain the route of the pipeline 2 with the best quality, and to find high-quality types, numbers, and combinations of the pipes 3 in the route. Furthermore, the route of the created pipeline 2 can be displayed on two drawings by the drawing output unit 4e. Therefore, for example, it becomes easier to give instructions when constructing the pipeline 2 on-site, or it becomes easier to share information about the pipeline 2 with others.
[0048] In addition, in the embodiment of the present invention, for the reward r in reinforcement learning, points are assigned in advance according to the separation distance from existing buried objects, the relative distance to the end point G, the arrival at the end point G, and the number of actions i. However, points may be assigned in advance according to at least one of the separation distance from existing buried objects, the relative distance to the end point G, the arrival at the end point G, and the number of actions i. In addition, in the embodiment of the present invention, as the characteristic values of the learning pipeline data 20 used when outputting the evaluation value of the learning pipeline data 20, the total length of the learning pipeline data 20, the average burial depth from the ground surface F of the learning pipeline data 20, and the number of times the learning pipeline data 20 passes below existing buried objects are used. However, at least one of the total length of the learning pipeline data 20, the average burial depth from the ground surface F of the learning pipeline data 20, and the number of times the learning pipeline data 20 passes below existing buried objects may be used.
[0049] In addition, in the embodiment of the present invention, in the selection pattern of the action a in reinforcement learning, there is only one type of straight pipe with a length of 1 m. However, it may be a selection pattern in which two or more types of straight pipes with different lengths can be selected, or the number of types of curved pipes with different curvatures may be set to 6 or less, or 8 or more types may be selectable. In addition, in the embodiment of the present invention, as shown in FIG. 7(b), the points to be assigned according to each environment and situation of the reward data 6b are determined. However, the points to be assigned may be other numbers. In addition, in the embodiment of the present invention, as the design route condition data 5B, setting parameters as shown in FIG. 6(a) can be set. However, other values of the design route conditions may be settable. Also, after storing the most highly evaluated learned pipeline data 20 output by the pipeline design system 1 of the present invention as the existing buried object data 5a in the underground design space 5A, the pipeline design system 1 may be operated again. Then, it is possible to output learned pipeline data 20 that bypasses the most highly evaluated learned pipeline data 20 output previously.
Industrial Applicability
[0050] The present invention is suitable for a pipeline design system that designs a pipeline capable of arranging electric wires or the like inside underground.
Explanation of Signs
[0051] 1…Pipeline design system 2…Pipeline 3…Pipe 3a…First selected pipe data 3b…Second selected pipe data 3c…Third selected pipe data 4…Information processing unit 4a…Design space setting unit 4b…Reinforcement learning unit 4c…Feature extraction unit 4d…Evaluation value output unit 4e…Drawing output unit 4f…Material quantity output unit 4g…Tension calculation unit 4h…Lateral pressure calculation unit 5…Input database 5A…Underground design space 5B…Design route condition data 5C…Learning condition data 5a…Existing buried object data 5b…Start point·end point position data 5c…Surface position data 5d…Region of minimum separation distance 5e…Region of required burial depth 6…Pipeline design database 6A…Reinforcement learning data 6B…Evaluation table 6a…Action data 6b…Reward data 7…Output database 7a…Planar pipeline data diagram 7b…Longitudinal section pipeline data diagram 7c…Connection material diagram data 7d…Material quantity data 7e…Tension calculation data 7f…Lateral pressure calculation data 7g…Route evaluation data 10…Computer 20…Learned pipeline data 20A…Learned pipeline data 20B…Learned pipeline data 20C…Learned pipeline data S…Start point G…End point F…Surface
Claims
1. A pipeline design system configured by sequentially connecting a plurality of pipes underground and capable of designing a pipeline that bypasses one or more existing buried objects, comprising: a design space setting unit that sets a three-dimensional underground design space; a reinforcement learning unit that selects the type of the pipe and its extending direction as actions, and based on a series of the actions when it is possible to move from a specified starting point to an ending point in the underground design space with the state where the data of each selected pipe connected in order is transitioned, learns and outputs one or more pieces of learned pipeline data with the highest total of preset rewards obtained respectively based on the actions; and is characterized by comprising the reinforcement learning unit.
2. In the pipeline design system according to Claim 1, the reward is a preset score according to at least one of the separation distance from the data of the existing buried object, the relative distance to the ending point, the reaching of the ending point, and the number of actions; and is characterized by the pipeline design system.
3. In the pipeline design system according to Claim 1 or 2, a feature extraction unit that calculates and extracts the feature values of a plurality of the learned pipeline data respectively; an evaluation value output unit that converts each of the feature values into an evaluation value based on a predetermined evaluation table and outputs the evaluation value respectively; and is characterized by comprising the evaluation value output unit.
4. In the pipeline design system according to Claim 3, the feature value is at least one of the total length of the learned pipeline data, the average buried depth from the ground surface of the learned pipeline data, and the number of times the learned pipeline data passes under the existing buried object; and is characterized by the pipeline design system.
5. In the pipeline design system according to Claim 1 or 2, a drawing output unit that outputs a planar pipeline data diagram when the underground design space in which the learned pipeline data is drawn is viewed from above and a longitudinal section pipeline data diagram when viewed from a predetermined side respectively; and is characterized by comprising the drawing output unit.
Citation Information
Patent Citations
A program for designing electrical utility trench conduits.
JP7086371B1