Detail level intelligent scheduling method for drawing digital twin three-dimensional scene of port and navigation
By designing reward functions in LOD technology and using ship trajectory prediction modules, the LOD data scheduling strategy is optimized, and the problem of mismatch between computing burden and pre-scheduling in the existing technology is solved, and the matching of resources and functional requirements and real-time drawing of dynamic scenarios is realized.
Patent Information
- Application Number
- CN202510073698.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-06-06
AI Technical Summary
The existing LOD technology has failed to comply with the functional requirements of different tasks in large-scale port and shipping digital twin 3D scenarios, and the dynamic three-dimensional scenario pre-scheduling has failed to match the functions of port and shipping digital twin system.
By designing reward functions, the LOD data scheduling strategy is optimized by combining the line of sight and importance, and adding the ship trajectory prediction module to generate the scene LOD data scheduling strategy and complete the pre-scheduling of the corresponding LOD data.
It realizes that computing resources are adapted to the functional requirements of different tasks in port and shipping digital twins, and dynamic three-dimensional scene LOD data pre-scheduling matches the functions of port and shipping digital twin systems, meeting different requirements for LOD data accuracy of different tasks.
Smart Images

Figure CN120107460A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional modeling, and in particular to a method for intelligent scheduling of detail levels of three-dimensional scene drawing of digital twins of ports and shipping lines. Background Art
[0002] In the rendering of large-scale digital twin scenes, the digital twin platform's ability to carry large-scale three-dimensional scene data can be significantly improved by simplifying the geometric mesh and constructing the Levels of Detail (LOD) of large-scale digital twin scenes. Among them, high-level LOD data is a simplification of low-level LOD data. In the implementation of LOD technology, the model uses different methods (such as polygon reduction, geometric mesh simplification, etc.) to generate model versions of different LOD levels. LOD data scheduling is usually based on the distance between the center of the model and the viewpoint (viewing distance), or based on the importance of the model in the scene. When drawing, the CPU will schedule the LOD level data required by the model to the GPU. Models that are farther away or less important will schedule low-level LOD data, while models that are close or more important will schedule high-level LOD data, thereby reducing unnecessary computing consumption in the GPU.
[0003] With the development of computer graphics technology, the models in modern virtual 3D scenes are becoming increasingly complex and detailed. At the same time, they are limited by limited hardware resources, which brings a huge computational burden to the rendering. Large-scale port and shipping 3D scenes include models of ships, buoys, terrain, quay cranes, cranes, buildings, green plants, etc. Millions to billions of polygons of these models need to be drawn in real time. In order to alleviate this large computational burden, LOD technology has made significant progress on the original basis. For example: LOD based on geometric simplification, LOD based on texture resolution, LOD based on volume and image, etc. These methods determine that the CPU can schedule different levels of LOD data in different scale scene requirements. Despite this, how to reduce the computational consumption in large-scale virtual scenes and improve the scene adaptability of LOD algorithms still needs to face some challenges.
[0004] At present, LOD technology still has some defects in large-scale port and shipping digital twin 3D scenes:
[0005] 1. The computational burden fails to adapt to the functional requirements of different tasks in the port and shipping digital twin. Traditional LOD technology usually schedules LOD data based on the view distance and importance of the model. When the model is close to the viewpoint or has a higher importance, the lower-level LOD data is scheduled, and when the view distance is far or the importance is low, the higher-level LOD data is scheduled. Although this scheduling mechanism can ensure that the computational burden is reduced to a certain extent, it ignores the different requirements for LOD data accuracy for different tasks in the three-dimensional scene of the port and shipping digital twin. For data that is not concerned with certain port and shipping tasks, the existing scheduling method may schedule over-fine LOD data, resulting in a waste of resources; and for data that is concerned with port and shipping tasks, the existing scheduling method may schedule coarse LOD data, resulting in failure to meet scene requirements. For example, in the task of ship entry, the LOD data of ships, waterways, berths, quay cranes and their surrounding infrastructure (such as lighthouses, buoys, etc.) should be scheduled at a finer level for drawing and spatial calculation to ensure that the ship's navigation system can accurately calculate the path, avoid obstacles and make real-time adjustments. The LOD levels of these port facilities need to be adjusted in a timely manner according to the progress of the mission to better schedule resources at the functional task level.
[0006] 2. The pre-scheduling of dynamic three-dimensional scenes fails to adapt to the functions of the port and shipping digital twin system. When the digital twin three-dimensional scene changes dynamically, the LOD data pre-scheduling technology will load the relevant LOD data into the video memory in advance so that it can be quickly retrieved during drawing to ensure the real-time drawing of dynamic scenes. Most of the existing LOD pre-scheduling technologies are based on different LOD data of the pre-scheduling model based on rules such as line of sight and importance. However, this pre-scheduling method ignores the priority and real-time drawing requirements of specific tasks in the three-dimensional scene of the port and shipping digital twin, resulting in the inability to match the drawing tasks with the functional tasks. For example, in the task of ship entry, when the ship is close to the dock, the channel drawing should pre-scheduling lower-level LOD data, while the traditional LOD pre-scheduling method often does not take into account the task scenario, resulting in the scheduling of the channel data LOD level The high-level LOD data appears, which does not meet the navigation task requirements of the ship in the digital twin scene. Summary of the invention
[0007] In view of the defects of the prior art, the present invention designs a reward function based on the viewing distance and importance of the three-dimensional port and shipping scenes, and optimizes the scene LOD data scheduling strategy through the advantage function calculated by the reward; adds a ship trajectory prediction module, determines the next scene state by predicting the ship coordinates, generates the scene LOD data scheduling strategy and completes the pre-scheduling of the corresponding LOD data, so as to ensure the real-time drawing of large-scale port and shipping twin scenes.
[0008] In order to achieve the above object, the present invention provides a method comprising the following steps:
[0009] (1) Build a three-dimensional digital twin scene of ports and shipping, set routes, and initialize environmental status data, which consists of three components: ship coordinates, scene LOD level, and importance;
[0010] (2) Construct prediction network, action network and evaluation network;
[0011] (3) Generate data and train prediction networks based on the three-dimensional scenes of port and shipping digital twins;
[0012] (3.1) At time t, the environmental state data s t Generate a scheduling action a through the action network t , in scheduling action a t Under the action of t+1 , and give a reward r(s t ,a t ); at the same time, s t The ship coordinate component h in t The prediction network generates the next scene ship coordinate h t+1 , with it and s t+1 Ship coordinate component x t+1 The mean square error between them is the loss function to optimize the prediction network;
[0013] (3.2) According to the set route, repeat step (3.1) to generate an experience trajectory data τ=(s 0 ,a 0 ,r 0 ,s 1 ,a 1 ,r 1 ,...) and stored in the data set; when the data set already contains N pieces of experience trajectory data, the newly generated data will replace the earliest experience trajectory data and be stored in the data set;
[0014] (3.3) Repeat the above operation multiple times until M pieces of empirical trajectory data are generated, M≤N;
[0015] (4) randomly sampling p pieces of experience trajectory data from the data set, P≤M, and sequentially inputting them into the evaluation network, calculating the advantage function based on the generated state value estimate, and optimizing the action network and the evaluation network using the advantage function;
[0016] (5) using the optimized action network, repeating steps (3) to (4) until the action network and the evaluation network converge or the number of training times is reached;
[0017] (6) The predicted coordinate values are obtained by using the trained prediction network, and the predicted coordinate values are used to replace the ship coordinates in the current environmental state data. The updated current environmental state data is input into the trained action network to generate a scheduling action, and the next environmental state is updated according to the scheduling action.
[0018] Furthermore, the reward is provided by the environment at each time step t, and is obtained based on the viewing distance and importance;
[0019] r final (a t ,s t )=a*r importance (a t ,s t )+b*r distance (a t ,s t )
[0020] Where: r distance (a t ,s t ) is the reward based on the viewing distance; r importance (a t ,s t ) is a reward based on importance; a and b are weighting coefficients and a+b=1.
[0021] Furthermore, the evaluation network and the action network are both neural networks, consisting of two fully connected layers, with 256 hidden layers and a Relu function as the activation function; the action network also needs to be connected to a normalization function Softmax at the end;
[0022] The prediction network adopts LSTM network.
[0023] Furthermore, the action network is specifically: the environment state s t Input the action network to generate a one-dimensional vector, which, after normalization, represents the sampling probability of different LOD levels; perform ε-greed sampling according to the probability, set the threshold ε, and when there is a probability greater than ε in the vector, take the action a corresponding to the maximum probability t As output; when the probability in the vector is less than ε, randomly sample action a according to the probability of the vector t ; In action a t Under the action of t+1 .
[0024] Furthermore, the step (4) is specifically as follows:
[0025] Select a group (s t ,s t+1) Input the evaluation network and output the state value estimate State value estimation and the reward r t Together they are used to calculate the advantage function
[0026]
[0027] Where: γ is a discount factor, γ∈(0,1);
[0028] Calculate the loss function of the action network based on the advantage function Loss Actor And the loss function Loss of the evaluation network Critic ;
[0029]
[0030] Where: ratio is the state s t Importance sampling obtained by input action network; clip is the truncation function. When the importance sampling exceeds the specified upper limit 1+τ or lower limit 1-τ, the function will return the corresponding upper or lower limit;
[0031] By minimizing the loss function Loss of the action network respectively Actor And the loss function Loss of the evaluation network Critic Update the action network and evaluation network;
[0032] Using the updated evaluation network, repeat the above operation until all randomly sampled p empirical trajectory data have been used.
[0033] Furthermore, the loss function of the prediction network is:
[0034] The input of the prediction network is the ship coordinate x at time t of the environmental state data t , the output is the predicted value h of the ship coordinates at time t+1 t+1 ;
[0035]
[0036] Among them: (M p ,H p ) is the predicted value of ship coordinates h t+1 The coordinates of (M T ,H T ) is s t+1 The ship coordinate component x in t+1 The coordinates of (M c ,H c ) are the coordinates on the set route.
[0037] The present invention also provides a detailed level intelligent scheduling system for port and shipping digital twin three-dimensional scene drawing, including:
[0038] The scene modeling module is used to build a three-dimensional scene of the port and shipping digital twin, set the route, and initialize the environmental status data, which consists of three components: ship coordinates, scene LOD level, and importance;
[0039] Network modeling module, used to build prediction network, action network and evaluation network;
[0040] Data generation and prediction network optimization module, used for data generation and prediction network training based on the three-dimensional scene of port and shipping digital twins;
[0041] (1) At time t, the environmental state data s t Generate a scheduling action a through the action network t , in scheduling action a t Under the action of t+1 , and give a reward r(s t ,a t ); at the same time, s t The ship coordinate component h in t The prediction network generates the next scene ship coordinate h t+1 , with it and s t+1 Ship coordinate component x t+1 The mean square error between them is the loss function to optimize the prediction network;
[0042] (2) According to the set route, repeat step (1) to generate an experience trajectory data τ=(s 0 ,a 0 ,r 0 ,s 1 ,a 1 ,r 1 ,...) and stored in the data set; when the data set already contains N pieces of experience trajectory data, the newly generated data will replace the earliest experience trajectory data and be stored in the data set;
[0043] (3) Repeat the above operation multiple times until M pieces of experience trajectory data are generated, M≤N;
[0044] An action network and evaluation network optimization module, used to randomly sample p pieces of experience trajectory data from the data set, P≤M, and sequentially input them into the evaluation network, calculate the advantage function according to the generated state value estimate, and optimize the action network and the evaluation network using the advantage function;
[0045] An updating module, used to use the optimized action network to iteratively complete training updates through a data generation and prediction network optimization module and an action network and evaluation network optimization module until the action network and the evaluation network converge or reach a training number;
[0046] The scheduling module is used to obtain the predicted coordinate value by using the trained prediction network, replace the ship coordinates in the current environmental state data with the predicted coordinate value, input the updated current environmental state data into the trained action network to generate a scheduling action, and realize the update of the next environmental state according to the scheduling action.
[0047] Beneficial effects of the present invention:
[0048] 1. Computing resources are adapted to the functional requirements of different tasks in the port and shipping digital twin.
[0049] The present invention designs rewards by comprehensively considering the viewing distance and importance, meeting the different requirements of LOD data accuracy for different tasks in the three-dimensional scene of the port and shipping digital twin. This enables the model in the scene to perform data scheduling with corresponding LOD accuracy according to the functional requirements of different tasks, and the computing resource allocation and functional task level are more matched.
[0050] 2. Dynamic three-dimensional scene LOD data pre-scheduling matches the functions of the port and shipping digital twin system.
[0051] The present invention predicts the coordinates of the ship in the next three-dimensional scene of the port digital twin through the ship trajectory prediction network, generates the LOD data scheduling strategy of the next scene using the environmental state formed by the predicted ship coordinates, and pre-schedules the corresponding LOD data. It meets the functional priority and real-time rendering requirements in the dynamic scene of the ship entering the port, making the rendering task more compatible with the functional task. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 A schematic flow chart of a detailed level intelligent scheduling method for drawing a three-dimensional scene of a digital twin of a port and shipping industry according to an embodiment of the present invention.
[0053] Figure 2 This is a structural schematic diagram of the detailed level intelligent scheduling method for drawing the three-dimensional scene of the digital twin of ports and shipping in an embodiment of the present invention.
[0054] Figure 3 Schematic diagram of the action network structure of an embodiment of the present invention.
[0055] Figure 4 This is a schematic diagram of the evaluation network structure of an embodiment of the present invention.
[0056] Figure 5 A design diagram of a reward mechanism based on a quadtree according to an embodiment of the present invention
[0057] Figure 6 This is a ship trajectory prediction point map according to an embodiment of the present invention. DETAILED DESCRIPTION
[0058] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0059] like Figure 1 As shown, an embodiment of the present invention provides a method for intelligent scheduling of detail levels of three-dimensional scene drawing of a port and shipping digital twin, including the following steps:
[0060] S101. Build a three-dimensional digital twin scene of ports and shipping, set routes, and initialize environmental status data. The environmental status data consists of three components: ship coordinates, scene LOD level, and importance.
[0061] The present invention uses the oblique photography technology of unmanned aerial vehicles to collect images of real port scenes; after the collection is completed, the processed images are reconstructed in three dimensions; the models in the reconstructed three-dimensional scene are singulated, and the entity objects in the three-dimensional scene (such as terrain, quay cranes, ships, cranes, buildings, green plants, etc.) are processed into independent three-dimensional models to realize the independent editing requirements of the models; and then the models are trimmed and edited respectively. For models with serious structural defects, the structure of the model is filled or replaced by 3dsmax modeling to ensure the integrity and sufficient details of the model. After the entire three-dimensional scene is deployed in Unity3D, the automatic LOD editing plug-in of Unity3D is used to generate four LOD-level model versions for each independent model with the same viewing distance and importance, so that the models in the three-dimensional scene can be displayed at different LOD levels. Among them, the high-level LOD data is a simplification of the low-level LOD data. Finally, a three-dimensional scene of digital twins of ports and shipping consisting of models such as terrain, quay cranes, ships, cranes, buildings, and green plants is formed.
[0062] Set the route, place the ship carrying the camera at the starting point of the route, and initialize the environmental state data at time t = 0. The environmental state data consists of three components: ship coordinates, scene LOD level, and importance. The ship coordinates are initialized to the starting point of the route, the LOD level is randomly initialized from the four levels (one is randomly selected), and the initialization of importance is classified according to the importance of the scene model.
[0063] S102: construct a prediction network, an action network and an evaluation network.
[0064] like Figure 2 As shown in the figure, a basic reinforcement learning framework for LOD data scheduling of port and shipping digital twin three-dimensional scenes is constructed, which consists of an action network and an evaluation network.
[0065] like Figure 3 , Figure 4As shown in the figure, both the evaluation network and the action network are neural networks, consisting of two fully connected layers, with 256 hidden layers, and the activation function is the rectified linear unit function (Relu). The action network is followed by a normalization function Softmax. The environmental state is used as the input of the action network (Actor Network) to generate the scheduling action of the scene LOD data. This data determines the scheduled LOD data for drawing new scenes. The evaluation network (Critic Network) is used for network training. The advantage function is calculated through the generated state value estimation, and the advantage function is used to optimize the action network and the evaluation network.
[0066] In the three-dimensional scene of the port and shipping digital twin, the data pre-scheduling of the next scene is an important technology to ensure the real-time drawing of the digital twin three-dimensional scene. The prediction network adopts the LSTM network, which takes the environmental state component - ship coordinates as the input of LSTM and outputs the predicted ship coordinates of the next scene; the environmental state updated with the predicted coordinate value generates the LOD scheduling strategy of the next scene, determines the LOD level of the next scene, and pre-schedules the corresponding LOD data into the video memory; when drawing the next scene, the renderer can quickly obtain data from the video memory, reducing the waiting time for data requests and data scheduling from memory to video memory.
[0067] The LSTM network workflow is represented by the following formula:
[0068] f t =σ(W f ·[h t-1 ,x t ]+b f )
[0069] i t =σ(W i ·[h t-1 ,x t ]+b i )
[0070] o t =σ(W o ·[h t-1 ,x t ]+b o )
[0071]
[0072] Where: W f , W i , W o and W c are the weight matrices for the forget gate, input gate, output gate, and updating the input matrix, respectively. f 、b i 、bo and b c are their bias, o t is the output matrix; σ is the sigmoid function, tanh is the hyperbolic tangent function, h t-1 is the predicted value of the ship coordinates at time t-1, [h t-1 , x t ] means concatenating the predicted value of the ship coordinates at time t-1 and the ship coordinates at time t into a longer vector.
[0073] S103. Generate data and train prediction networks based on the three-dimensional scene of port and shipping digital twins.
[0074] (1) At time t, the environmental state data s t Generate scheduling action a through action network t , in scheduling action a t Under the action of t+1 , and give a reward r(s t ,a t ); at the same time, s t The ship coordinate component h in t Generate the next scene ship coordinate h through the prediction network t+1 , with it and s t+1 Ship coordinate component x t+1 The mean square error between them is the loss function to optimize the prediction network.
[0075] At time t, the environment state s t Input the action network and output a one-dimensional vector. After normalization, it represents the sampling probability of different LOD levels. Perform ε-greed sampling according to the probability, set the threshold ε, and when there is a probability greater than ε in the vector, take the action a corresponding to the maximum probability t As output; when the probability in the vector is less than ε, randomly sample action a according to the probability of the vector t ; In action a t Under the action of t+1 And get reward r(s t ,a t ).
[0076] In the embodiment of the present invention, the LOD level is divided into 4 levels, and ε is 0.5.
[0077] Reward r(s t ,a t ) is provided by the environment at each time step t to represent the state s t Next, perform action a tThe quality of the LOD data directly affects the scheduling of LOD data of the three-dimensional scene of the port and shipping digital twin. The embodiment of the present invention comprehensively considers the viewing distance and importance to design the scene LOD data scheduling reward mechanism.
[0078] The reward mechanism based on viewing distance is designed by quadtree principle to calculate the reward r distance (a t ,s t ).like Figure 5 As shown in the figure, the scene in the viewport is divided into grids. The terminal node of the quadtree stores scene data with a grid size of 64*64 (unit: m), which represents scene LOD1; the grid size is downsampled to 2*2, and the scene data with grid sizes of 128*128 and 256*256 are stored in the leaf nodes, which represent scene LOD2 data and scene LOD3 data respectively; the root node describes the scene with a grid size of 512*512, which represents scene LOD4 data. So far, the scene in the viewport is divided into four levels from LOD1 to LOD4 from near to far. When calculating the reward, the scene LOD data is scheduled according to the distance between the center of the scene model and the viewpoint (viewing distance). Based on the diagonal length of the grid size in different LOD levels, the visible distance is divided into four intervals along the grid diagonal: [0,181], [181,362], [362,724], [724,1448] (unit: m). The LOD data scheduling strategy determines the LOD level of the scene. The reward rules are designed according to different view distances as follows: when the view distance is within [0, 181] and the strategy is LOD1, the reward is 1, otherwise it is 0; when the view distance is within [181, 362] and the strategy is LOD2 or LOD3, the reward is 1, otherwise it is 0; when the view distance is within [362, 724] and the strategy is LOD3 or LOD4, the reward is 1, otherwise it is 0; when the view distance is within [724, 1448] and the strategy is LOD4, the reward is 1, otherwise it is 0. The reward r based on the view distance is calculated by the above rules. distance (a t ,s t ), the calculation formula is as follows:
[0079]
[0080] Where d represents the viewing distance.
[0081] Calculate reward r based on importance importance (a t ,s t ), the models in the scene are divided according to the importance of different scenes. According to the knowledge of ship navigation, the models in the scene are divided into two categories according to their importance. The first category is Crucial1, which includes buoys, quay cranes, cranes, container trucks and ships; the second category is Crucial2, and the remaining models except the first category are grouped together. Rewards r are calculated based on importanceimportance (a t ,s t ) is as follows: if the scene model belongs to Crucial1 and the scheduling strategy is LOD1 or LOD2, the reward is 1, otherwise it is 0; if the scene model belongs to Crucial2 and the scheduling strategy is not lower than LOD3, the reward is 1, otherwise it is 0. importance (a t ,s t ) is calculated as follows:
[0082]
[0083] Among them, imp refers to the importance of the scene.
[0084] Finally, the two rewards r are assigned according to weights a and b. distance (a t ,s t ) and r importance (a t ,s t ) is weighted summed as the final reward r for scene LOD data scheduling final (a t ,s t )
[0085] r final (a t ,s t )=a*r importance (a t ,s t )+b*r distance (a t ,s t )
[0086] Among them: a+b=1.
[0087] Predicting network usage environment status t The predicted coordinate value of the ship coordinate component in the image is obtained and compared with the new environment state s t+1 The mean square error between the coordinate components of the ship in is used as the objective function to train the prediction network. Since ships need to comply with the requirements of lane separation and designed routes during navigation, based on this premise, the embodiment of the present invention designs a distance loss function L loc ,like Figure 6 As shown, the triangle point is the predicted value of the ship coordinate h t+1 Coordinates (M p ,H p ), the square point is the coordinate on the set route (M c ,H c ), the circular point is s t+1 The ship coordinate component x int+1 Coordinates (M T ,H T ). When the predicted coordinates are outside the designed route, (M p -M T ,H p -H T ) and (M c -M T ,H c -H T ) The angle β between the two vectors is greater than 90 degrees; when the predicted coordinate is inside the designed route, the angle β between the two vectors is less than 90 degrees, and the calculation formula of β is as follows:
[0088]
[0089] Different strategies are implemented for the predicted points that deviate from the ship's design route and the predicted points that do not deviate from the design route. When 0≤β≤90, the loss function calculation formula is set to the L2 distance. When 90≤β≤180, this paper adds a penalty coefficient (1-cosβ) to the L2 loss value. The purpose is to give a larger loss value to the predicted value that deviates from the design route and the outside of the actual trajectory point to help the model converge quickly in a direction. The loss function calculation formula is as follows:
[0090]
[0091] The prediction network is trained with the goal of minimizing the above loss function.
[0092] (2) According to the ship's route setting, repeat step (1) to generate an experience trajectory data consisting of state, action, and reward τ = (s 0 ,a 0 ,r 0 ,s 1 ,a 1 ,r 1 ,...) are stored in the data set; when the data set already contains N pieces of experience trajectory data, the newly generated data will replace the earliest experience trajectory data and be stored in the data set.
[0093] In the embodiment of the present invention, N=5000, the empirical trajectory data τ=(s 0 ,a 0 ,r 0 ,s 1 ,a 1 ,r 1 ,...) includes data at each time point from the starting point to the end point of the route.
[0094] (3) Repeat the operation multiple times until M pieces of experience trajectory data are generated, M ≤ N.
[0095] Multiple experiments generate batches of experience trajectory data and store them in the experience buffer pool to improve the utilization of historical data.
[0096] In the embodiment of the present invention, M=1000.
[0097] S104, randomly sampling p pieces of experience trajectory data from all the data in the data set, P≤M, and sequentially inputting them into the evaluation network, calculating the advantage function according to the generated state value estimate, and optimizing the action network and the evaluation network using the advantage function.
[0098] Select a group (s t ,s t+1 ) Input evaluation network, output state value estimate State value estimation and the reward r t Together they are used to calculate the advantage function
[0099]
[0100] Where: γ is a discount factor, γ∈(0,1).
[0101] Calculate the loss function of the action network based on the advantage function Loss Actor And the loss function Loss of the evaluation network Critic .
[0102]
[0103]
[0104] Where: ratio is the state s t Importance sampling obtained by inputting the action network; clip is the truncation function. When the importance sampling exceeds the specified upper limit 1+τ or lower limit 1-τ, the function will return the corresponding upper or lower limit.
[0105] By minimizing the loss function Loss of the action network respectively Actor And the loss function Loss of the evaluation network Critic Update the action network and evaluation network.
[0106] Using the updated evaluation network, repeat the above operation until all randomly sampled p empirical trajectory data have been used.
[0107] In the embodiment of the present invention, p=64, τ=0.2, and γ=0.9.
[0108] S105. Using the optimized action network, repeat steps S103-S104 until the action network and the evaluation network converge or the maximum number of training times is reached.
[0109] S106, using the trained prediction network to obtain predicted coordinate values, replacing the ship coordinates in the current environmental state data with the predicted coordinate values, inputting the updated current environmental state data into the trained action network to generate a scheduling action, and implementing the update of the next environmental state according to the scheduling action.
[0110] After training, the prediction network can accurately predict the coordinates of the ship in the next scene, and the action network also has a scene LOD data scheduling strategy that is suitable for port and shipping tasks. During testing, the environment state is initialized, and the environment state is input into the action network and trajectory prediction network at the same time. The action network generates the LOD data scheduling strategy, and the scene starts drawing using the scheduled data. The prediction network generates the coordinates of the ship in the next scene, which are used to replace the ship coordinates in the current environment state data. The updated environment state is input into the action network to generate the LOD scheduling strategy for the next scene, and the LOD data corresponding to the strategy is scheduled from the memory to the video memory. After the scene is drawn, the video memory has stored the LOD data required to draw the next scene. When the ship moves and switches to the next scene, the LOD data is directly retrieved from the video memory to start the scene drawing, and the scene state generates the LOD scheduling strategy through the action network for data pre-scheduling.
[0111] The present invention also provides a detailed level intelligent scheduling system for port and shipping digital twin three-dimensional scene drawing, including:
[0112] The scene modeling module is used to build a three-dimensional scene of the port and shipping digital twin, set the route, and initialize the environmental status data, which consists of three components: ship coordinates, scene LOD level, and importance;
[0113] Network modeling module, used to build prediction network, action network and evaluation network;
[0114] Data generation and prediction network optimization module, used for data generation and prediction network training based on the three-dimensional scene of port and shipping digital twins;
[0115] (1) At time t, the environmental state data s t Generate a scheduling action a through the action network t , in scheduling action a t Under the action of t+1 , and give a reward r(s t ,a t ); at the same time, s t The ship coordinate component h in tThe prediction network generates the next scene ship coordinate h t+1 , with it and s t+1 Ship coordinate component x t+1 The mean square error between them is the loss function to optimize the prediction network;
[0116] (2) According to the set route, repeat step (1) to generate an experience trajectory data τ=(s 0 ,a 0 ,r 0 ,s 1 ,a 1 ,r 1 ,...) and stored in the data set; when the data set already contains N pieces of experience trajectory data, the newly generated data will replace the earliest experience trajectory data and be stored in the data set;
[0117] (3) Repeat the above operation multiple times until M pieces of experience trajectory data are generated, M≤N;
[0118] An action network and evaluation network optimization module, used to randomly sample p pieces of experience trajectory data from the data set, P≤M, and sequentially input them into the evaluation network, calculate the advantage function according to the generated state value estimate, and optimize the action network and the evaluation network using the advantage function;
[0119] An updating module, used to use the optimized action network to iteratively complete training updates through a data generation and prediction network optimization module and an action network and evaluation network optimization module until the action network and the evaluation network converge or reach a training number;
[0120] The scheduling module is used to obtain the predicted coordinate value by using the trained prediction network, replace the ship coordinates in the current environmental state data with the predicted coordinate value, input the updated current environmental state data into the trained action network to generate a scheduling action, and realize the update of the next environmental state according to the scheduling action.
[0121] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the principles and spirit of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for intelligent scheduling of detail levels for three-dimensional scene drawing of digital twins of ports and shipping, characterized in that: The steps include: (1) Build a three-dimensional digital twin scene of ports and shipping, set routes, and initialize environmental status data, which consists of three components: ship coordinates, scene LOD level, and importance; (2) Construct prediction network, action network and evaluation network; (3) Generate data and train prediction networks based on the three-dimensional scene of port and shipping digital twins; (3.1) At time t, the environmental state data s t Generate a scheduling action a through the action network t , in scheduling action a t Under the action of t+1 , and give a reward r(s t ,a t ); at the same time, s t The ship coordinate component h in t The prediction network generates the next scene ship coordinate h t+1 , with it and s t+1 Ship coordinate component x t+1 The mean square error between them is the loss function to optimize the prediction network; (3.2) According to the set route, repeat step (3.1) to generate an experience trajectory data τ = (s0, a0, r0, s1, a1, r1, ...) consisting of state, action, and reward from the starting point to the end point of the set route and store it in the data set; when the data set already contains N experience trajectory data, the newly generated data will replace the earliest experience trajectory data and be stored in the data set; (3.3) Repeat the above operation several times until M pieces of empirical trajectory data are generated, M≤N; (4) randomly sampling p pieces of experience trajectory data from the data set, P≤M, and sequentially inputting them into the evaluation network, calculating the advantage function based on the generated state value estimate, and optimizing the action network and the evaluation network using the advantage function; (5) using the optimized action network, repeating steps (3) to (4) until the action network and the evaluation network converge or the number of training times is reached; (6) The predicted coordinate values are obtained by using the trained prediction network, and the predicted coordinate values are used to replace the ship coordinates in the current environmental state data. The updated current environmental state data is input into the trained action network to generate a scheduling action, and the next environmental state is updated according to the scheduling action.
2. The method for intelligent scheduling of detail levels for three-dimensional scene rendering of port and shipping digital twins according to claim 1 is characterized by: The reward is provided by the environment at each time step t, based on the viewing distance and importance. r final (a t ,s t )=a*r importance (a t ,s t )+b*r distance (a t ,s t ) Where: r distance (a t ,s t ) is the reward based on the viewing distance; r importance (a t ,s t ) is a reward based on importance; a and b are weighting coefficients and a+b=1.
3. The method for intelligent scheduling of detail levels for three-dimensional scene rendering of port and shipping digital twins according to claim 1 is characterized by: The evaluation network and the action network are both neural networks, consisting of two fully connected layers, with 256 hidden layers and a Relu function as the activation function; the action network also needs to be connected to a normalization function Softmax at the end; The prediction network adopts LSTM network.
4. The method for intelligent scheduling of detail levels for three-dimensional scene rendering of port and shipping digital twins according to claim 1 is characterized in that: The action network is specifically: the environment state s t Input the action network to generate a one-dimensional vector, which, after normalization, represents the sampling probability of different LOD levels; perform ε-greed sampling according to the probability, set the threshold ε, and when there is a probability greater than ε in the vector, take the action a corresponding to the maximum probability t As output; when the probability in the vector is less than ε, randomly sample action a according to the probability of the vector t ; In action a t Under the action of t+1 .
5. The method for intelligent scheduling of detail levels for three-dimensional scene rendering of port and shipping digital twins according to claim 1 is characterized in that: The step (4) is specifically: Select a group (s t ,s t+1 ) Input the evaluation network and output the state value estimate V t * , State value estimation V t * , and the reward r t Together they are used to calculate the advantage function Where: γ is a discount factor, γ∈(0,1); Calculate the loss function of the action network based on the advantage function Loss Actor And the loss function Loss of the evaluation network Critic ; Where: ratio is the state s t Importance sampling obtained by input action network; clip is the truncation function. When the importance sampling exceeds the specified upper limit 1+τ or lower limit 1-τ, the function will return the corresponding upper or lower limit; By minimizing the loss function Loss of the action network respectively Actor And the loss function Loss of the evaluation network Critic Update the action network and evaluation network; Using the updated evaluation network, repeat the above operation until all randomly sampled p empirical trajectory data have been used.
6. The method for intelligent scheduling of detail levels for three-dimensional scene rendering of port and shipping digital twins according to claim 1 is characterized in that: The loss function of the prediction network is: The input of the prediction network is the ship coordinate x at time t of the environmental state data t , the output is the predicted value h of the ship coordinates at time t+1 t+1 Among them: (M p ,H p ) is the predicted value of ship coordinates h t+1 The coordinates of (M T ,H T ) is s t+1 The ship coordinate component x in t+1 The coordinates of (M c ,H c ) are the coordinates on the set route.
7. A detailed level intelligent scheduling system for port and shipping digital twin three-dimensional scene drawing, characterized by: include: The scene modeling module is used to build a three-dimensional scene of the port and shipping digital twin, set the route, and initialize the environmental status data, which consists of three components: ship coordinates, scene LOD level, and importance; Network modeling module, used to build prediction network, action network and evaluation network; Data generation and prediction network optimization module, used for data generation and prediction network training based on the three-dimensional scene of port and shipping digital twins; (1) At time t, the environmental state data s t Generate a scheduling action a through the action network t , in scheduling action a t Under the action of t+1 , and give a reward r(s t ,a t ); at the same time, s t The ship coordinate component h in t The prediction network generates the next scene ship coordinate h t+1 , with it and s t+1 Ship coordinate component x t+1 The mean square error between them is the loss function to optimize the prediction network; (2) According to the set route, repeat step (1) to generate an experience trajectory data τ = (s0, a0, r0, s1, a1, r1, ...) consisting of state, action, and reward from the starting point to the end point of the set route and store it in the data set; when the data set already contains N experience trajectory data, the newly generated data will replace the earliest experience trajectory data and be stored in the data set; (3) Repeat the above operation multiple times until M pieces of experience trajectory data are generated, M≤N; An action network and evaluation network optimization module, used to randomly sample p pieces of experience trajectory data from the data set, P≤M, and sequentially input them into the evaluation network, calculate the advantage function according to the generated state value estimate, and optimize the action network and the evaluation network using the advantage function; An updating module, used to use the optimized action network to iteratively complete training updates through a data generation and prediction network optimization module and an action network and evaluation network optimization module until the action network and the evaluation network converge or reach a training number; The scheduling module is used to obtain the predicted coordinate value by using the trained prediction network, replace the ship coordinates in the current environmental state data with the predicted coordinate value, input the updated current environmental state data into the trained action network to generate a scheduling action, and realize the update of the next environmental state according to the scheduling action.