Virtual reality navigation method

By constructing 3D scenes and using acoustic scanning data to model the terrain in real time, the problem that traditional underwater navigation technology is difficult to provide an intuitive and comprehensive navigation experience is solved, and efficient restoration and real-time navigation of underwater operation scenarios are achieved, improving the safety and efficiency of operations.

CN120163948APending Publication Date: 2025-06-17Shanghai Salvage Bureau of the Ministry of Transport
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510195069.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Traditional underwater navigation technology is difficult to provide an intuitive and comprehensive navigation experience, especially in scenarios with high requirements for complex terrain or posture.

Method used

By constructing 3D scenes, using acoustic scanning data to model terrain in real time, combining the real-time data placement, movement trajectory and pose of divers or submersibles, it provides multi-view switching and intelligent target recognition functions.

Benefits of technology

It realizes efficient restoration and real-time navigation of underwater operation scenarios, improves the safety and efficiency of operations, and provides efficient and accurate navigation solutions in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163948A_ABST
    Figure CN120163948A_ABST
Patent Text Reader

Abstract

The invention provides a virtual reality navigation method, which belongs to the technical field of navigation, and comprises the following specific steps: S1, constructing a 3D scene; s2, loading a terrain model in the 3D scene, and displaying data such as the position, the action track and the posture of the diver or the submersible vehicle; s3, performing incremental updating on the terrain model according to real-time data of acoustic scanning; s4, the display position, the action track and the posture of the diver or the submersible vehicle are updated according to the real-time data pushed by the diver or the submersible vehicle; and S5, performing diving command, and performing judgment according to data such as the display position, the action track and the posture of the diver or the submersible vehicle so as to guide the diver or the submersible vehicle to perform operation. According to the invention, the underwater operation can be safer and more efficient. And meanwhile, an efficient and accurate underwater navigation solution is provided in a complex environment and severe weather, so that the working efficiency is improved and errors are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of navigation and relates to a virtual reality navigation method. Background Art

[0002] Deep-sea diving operations have poor visibility, high environmental risks, and difficult communication, resulting in low operation efficiency. Long-term operations pose a double test to the physiology and psychology of divers. Traditional underwater navigation technologies are limited to two-dimensional planes and simple data interactions, and it is difficult to provide an intuitive and comprehensive navigation experience, especially in complex terrain or scenarios with high attitude requirements. Summary of the Invention

[0003] To solve the above technical problems, the present invention provides a virtual reality navigation method, which constructs a 3D scene to restore the underwater operation scene one by one. Through the three-dimensional reproduction of the operation scene in multiple dimensions, the operation command can comprehensively and intuitively view the global scene of the diving operation, and the operation states such as the position, action trajectory, and attitude of the diver in the scene are presented in real time. The terrain model in the scene is completed by real-time modeling of the seabed terrain using acoustic survey data, and through morphological restoration technology, the underwater 3D scene is made more vivid and realistic.

[0004] The technical solution adopted by the present invention is as follows: A virtual reality navigation method, the specific steps are as follows: S1, construct a 3D scene; S2, load the terrain model, as well as the data of the display position, action trajectory, and attitude of the diver or submersible in the 3D scene; S3, the terrain model is incrementally updated according to the real-time data of acoustic survey; S4, the display position, action trajectory, and attitude of the diver or submersible are updated according to the real-time data pushed by the diver or submersible; S5, the diving command makes a judgment based on the data such as the display position, action trajectory, and attitude of the diver or submersible to guide the operation.

[0005] Further, the construction of the three-dimensional scene in step S1 is to use the JavaScript framework Vue.js to build the front-end architecture of the system, and use the JavaScript framework Three.js to build the 3D scene.

[0006] Further, the modeling steps of the terrain model in step S2 are as follows: S21, obtain environmental information data through refined multi-beam survey; S22, generate three-dimensional point clouds from the environmental information data in step S21 and process them using libraries such as PCL (Point Cloud Library); S23. Apply a filtering algorithm (voxel grid filter) to reduce the point cloud density and simplify data processing, which can remove useless data. Use the feature calculation module provided by PCL to extract point cloud features (such as normal vectors, curvatures, etc.) for subsequent processing; S24. Use algorithms (such as Poisson Surface Reconstruction, Greedy Triangulation) to convert the point cloud features into a 3D mesh model; S25. Use PCL or Open3D to apply color information to the generated mesh and use UV mapping technology to paste the image onto the surface of the 3D model.

[0007] Furthermore, the terrain model adds a rainbow map mode, and the specific implementation steps are as follows: S201. Obtain terrain elevation data; S202. Calculate the overall height of the terrain, and set a color demarcation point using the ratio of the current height to the overall height; The formula is as follows: maxy = yArr.reduce((a, b) => (a > b? a : b)); miny = yArr.reduce((a, b) => (a < b? a : b)); height = maxy - miny; percent = (pos.getZ(i) - miny) / height; c1 = color1.clone().lerp(color2, percent * 1.5); c2 = color2.clone().lerp(color3, (percent - 0.32) * 1.7); Among them, maxy represents the maximum value of elevation data, a represents the current comparison data, b represents the next comparison data, miny represents the minimum value of elevation data, yArr represents the set of elevation data, height represents the overall height of the mountain range, percent represents the ratio of the current height to the overall height and the number of iterations of the i loop, pos represents the set of terrain vertex data, color1 represents the valley color, color2 represents the mountainside color, color3 represents the mountaintop color, c1 represents the interpolation between color1 and color2, c2 represents the interpolation between color2 and color3, reduce can iteratively compare the data in the elevation data set to find the maximum and minimum values, the getZ method can filter out the Z-axis data of the coordinates, the clone method is to copy an object, and lerp can regenerate the color value through linear interpolation.

[0008] S203. Set the color attribute of the color of the model data attributes in the form of color interpolation.

[0009] Furthermore, the incremental update steps of the terrain model in step S3 are as follows: S31. Perform duplicate checking on the newly scanned real-time data; S32. Merge the deduplicated real-time data into the existing terrain model; S33. After incrementally updating the existing terrain model, switch it in the scene in real time.

[0010] Furthermore, the movement trajectory of the diver or submersible in step S2 is created using the built-in Line object of Three.js of the JavaScript framework. First, create a BufferGeometry (geometry) and a LineBasicMaterial (line material), and then create a Line line object and add the geometry and material to the object.

[0011] Furthermore, the update formula of the movement trajectory in step S4 is as follows: positions[i×3]=positionHistory[i].x positions[i×3+1]=positionHistory[i].y positions[i×3+2]=positionHistory[i].z maxPoints = positionHistory[i] + maxPoints where positionHistory represents an array of storage points, maxPoints represents the maximum number, positions represents vertex data, i represents the loop iteration count, positionHistory[i].x represents obtaining the x-axis data of the current point, positionHistory[i].y represents obtaining the y-axis data of the current point, and positionHistory[i].z represents obtaining the z-axis data of the current point.

[0012] Furthermore, the attitude of the diver or submersible in step S2 is simulated by rotating the model position. It is necessary to calculate the center point of the model as the rotation center. The calculation formula for the center point is as follows: boundingBox = Box3(min, max) centerPosition = (boundingBox.min + boundingBox.max / 2,) where boundingBox is the bounding box of the model calculated by THREE.Box3().setFromObject(data) (the bounding box of the three-dimensional axis alignment), min represents the minimum coordinate, max represents the maximum value, BOX3 is the built-in object of the three.js model bounding box, and centerPosition represents the center point.

[0013] Furthermore, the real-time data pushed by the diver or submersible in step S4 is the real-time state of the current navigation environment and real-time depth information obtained by collecting relevant data with wave and current sensors, tide sensors, and light sensors and integrating the data from different sensors using data fusion technology.

[0014] Furthermore, the real-time data pushed by the diver or submersible in step S4 also includes images. The images can be intelligently recognized and located for various underwater objects through the deep learning object detection framework YOLO. The specific steps are as follows: S41, YOLO regards the object detection task as a regression problem and directly predicts the bounding box and class probability from the image pixels; S42, divide the input image into a grid of S×S. Each grid is responsible for predicting a fixed number of bounding boxes and their confidence levels and class probabilities; S43, the bounding box predicted by each grid consists of four coordinates (center point position, width, height) and a confidence value. The confidence represents the probability that the box contains an object; S44, use the Softmax function to predict the classes for each grid, which can identify multiple classes of objects simultaneously.

[0015] Adopting the technical solution of the present invention can achieve the following beneficial effects: Functions such as underwater operation scenario restoration, real-time path navigation in 3D scenes, display of action trajectories and postures, multi-view switching, real-time seabed modeling, rainbow maps, environmental perception, and intelligent target recognition are realized, making underwater operations safer and more efficient. At the same time, in complex environments and bad weather, an efficient and accurate underwater navigation solution is provided, which helps to improve operation efficiency and reduce errors. Brief Description of the Drawings

[0016] The following briefly describes the content expressed in each drawing of this specification and the marks in the drawings: Figure 1 It is a schematic structural diagram of the 3D scene of the present invention; Figure 2 It is a schematic diagram after adding the rainbow map mode to the terrain model of the present invention; Figure 3 It is a schematic diagram of terrain model modeling of the present invention; Figure 4 It is a schematic diagram of the terrain model updated after surveying of the present invention; Figure 5 It is a reference diagram of the effect after environmental data fusion processing of the present invention; Figure 6 It is a schematic diagram of intelligent image recognition and classification of objects of the present invention. Detailed Embodiment

[0017] The following further details the specific embodiments of the present invention, such as the shapes, structures of the components involved, the mutual positions and connection relationships between the parts, the functions of the parts, and the working principles, etc., by describing the embodiments with reference to the drawings.

[0018] This embodiment provides a virtual reality navigation method, and its specific steps are as follows: S1, construct a 3D scene; the construction of the 3D scene is to use the JavaScript framework Vue.js to construct the front-end architecture of the system, and use the JavaScript framework Three.js to construct the 3D scene, as Figure 1 shown.

[0019] S2, load the terrain model, as well as the display positions, action trajectories, and posture data of divers or submersibles in the 3D scene; S3, the terrain model is incrementally updated according to the real-time data of acoustic surveying; S4, the display positions, action trajectories, and postures of divers or submersibles are updated according to the real-time data pushed by divers or submersibles; In S5, the diving commander makes judgments based on data such as the displayed position, movement trajectory, and attitude of the diver or submersible to guide the operation.

[0020] The three-dimensional scene construction in step S1 uses the JavaScript framework Vue.js to build the front-end architecture of the system and the JavaScript framework Three.js to build the 3D scene.

[0021] The modeling steps of the terrain model in step S2 are as follows: S21, obtain environmental information data through refined multi-beam sounding; S22, generate a three-dimensional point cloud from the environmental information data in step S21 and process it using libraries such as PCL (Point Cloud Library); S23, apply a filtering algorithm (voxel grid filter) to reduce the point cloud density and simplify data processing, which can remove useless data, and use the feature calculation module provided by PCL to extract point cloud features (such as normal vectors, curvatures, etc.) for subsequent processing; S24, use algorithms (such as Poisson Surface Reconstruction, Greedy Triangulation) to convert the point cloud features into a three-dimensional grid model; S25, use PCL or Open3D to apply color information to the generated grid and use UV mapping technology to paste the image onto the surface of the three-dimensional model, as shown in Figure 3 shown.

[0022] Since the terrain model has undulations, in order to facilitate viewing of the terrain changes and vividly display the change trend of the terrain elevation, the terrain model adds a rainbow map mode, and the effect is as shown in Figure 2 shown. The specific implementation steps are as follows: S201, obtain terrain elevation data; S202, calculate the overall height of the terrain, use the ratio of the current height to the overall height to set a color demarcation point; The formula is as follows: maxy = yArr.reduce((a, b) => (a > b? a : b)); miny = yArr.reduce((a, b) => (a < b? a : b)); height = maxy - miny; percent = (pos.getZ(i) - miny) / height; c1 = color1.clone().lerp(color2, percent * 1.5); c2 = color2.clone().lerp(color3, (percent - 0.32) * 1.7); Among them, maxy represents the maximum value of elevation data, a represents the current comparison data, b represents the next comparison data, miny represents the minimum value of elevation data, yArr represents the set of elevation data, height represents the overall height of the mountain range, percent represents the ratio of the current height to the overall height and the number of iterations of the i loop, pos represents the set of terrain vertex data, color1 represents the valley color, color2 represents the mountainside color, color3 represents the mountaintop color, c1 represents the interpolation between color1 and color2, c2 represents the interpolation between color2 and color3, reduce can iteratively compare the data in the elevation data set to find the maximum and minimum values, the getZ method can filter out the Z-axis data of the coordinates, the clone method is to copy an object, and lerp can regenerate the color value through linear interpolation.

[0023] S203. Set the color attribute of the color of the model data attributes in the form of color interpolation.

[0024] In step S2, the movement trajectory of the diver or submersible is created using the built-in Line object of Three.js of the JavaScript framework. First, create BufferGeometry (geometry) and LineBasicMaterial (line material), and then create a Line line object and add the geometry and material to the object.

[0025] In step S2, the attitude of the diver or submersible is simulated by rotating the model position. It is necessary to calculate the center point of the model as the rotation center. The calculation formula for the center point is as follows: boundingBox = Box3(min, max) centerPosition = (boundingBox.min + boundingBox.max / 2,) Among them, boundingBox is the bounding box of the model calculated by THREE.Box3().setFromObject(data) (the boundary box of the three-dimensional axis alignment), min represents the minimum coordinate value, max represents the maximum value, BOX3 is the built-in object of the three.js model bounding box, and centerPosition represents the center point.

[0026] The OrbitControls (controller) is used for the rotation, zooming, and movement control of the free view. The Camera provides different view switching functions, supporting free view, top view, and side view. Through geometric operations, the distance between the diver and two points, namely the target marker point or the terrain, is calculated in real time. The formula is as follows: distance = pointA.distanceTo(pointB) pointA represents the diver's position, pointB represents the target position, and the distanceTo method is a built-in method in threejs that can calculate the distance between two three-dimensional coordinates.

[0027] S3. The terrain model is incrementally updated according to the real-time data of the acoustic survey; incremental updates are made after each update of the survey data, that is, the details of the existing model are retained to enhance the richness of details.

[0028] The steps for the incremental update of the terrain model are as follows: S31. Duplicate-check the newly surveyed real-time data; ensure that only new data is added; S32. Merge the deduplicated real-time data into the existing terrain model; use a smoothing algorithm (such as Laplacian smoothing) to process the boundaries and reduce visual discontinuities; S33. After the incremental update of the existing terrain model, switch it in the scene in real time. Make an incremental update to the original model, keep the overall structure and texture of the model consistent, export the new model, and notify the update through websocket. The scene automatically replaces the new model, as shown in Figure 4 shown.

[0029] S4. The display position, movement trajectory, and attitude of the diver or submersible are updated according to the real-time data pushed by the diver or submersible, and the movement trajectory is pushed to the diver or submersible in real time; The update formula for the movement trajectory is as follows: positions[i×3]=positionHistory[i].x positions[i×3+1]=positionHistory[i].y positions[i×3+2]=positionHistory[i].z maxPoints = positionHistory[i] + maxPoints Among them, positionHistory represents the array of stored points, maxPoints represents the maximum number, positions represents the vertex data, i represents the number of loop iterations, positionHistory[i].x represents getting the x-axis data of the current point, positionHistory[i].y represents getting the y-axis data of the current point, and positionHistory[i].z represents getting the z-axis data of the current point.

[0030] The real-time data pushed by divers or submersibles is collected by carrying wave and water flow sensors, tidal sensors and optical sensors, and the data from different sensors are integrated by data fusion technology to obtain the real-time status and real-time depth information of the current navigation environment. Specifically, flow sensors and wave sensors are installed to monitor water flow speed and wave height in real time. These sensors can work on ultrasonic or electromagnetic principles to provide accurate environmental data. Use tidal sensors to monitor water level changes to obtain information about tidal cycles. This can be achieved by integrating air pressure sensors. Optical sensors are installed to detect changes in the sun's angle and underwater lighting to provide reference information about diving depth and time. Through the above data collection, data fusion technology is used to integrate data from different sensors to provide more comprehensive environmental status information. For example, combine water flow, wave and tidal data to calculate the real-time status of the current navigation environment. Combine the data of the depth detector with the environmental data to provide real-time depth information for divers or submersibles. This can help the diving commander determine when to adjust the diving direction, avoid potential dangers, etc., to achieve the effect reference Figure 5 .

[0031] The real-time data pushed by divers or submersibles also includes images, which can be intelligently identified and located by YOLO, a deep learning object detection framework. The specific steps are as follows: S41, YOLO treats the object detection task as a regression problem, directly predicting bounding boxes and class probabilities from image pixels; S42, divides the input image into S×S grids, each grid is responsible for predicting a fixed number of bounding boxes and their confidence and category probabilities; S43, the bounding box predicted by each grid consists of four coordinates (center point position, width, height) and a confidence value, where the confidence value indicates the probability that the box contains an object; S44, using the Softmax function to predict the category of each grid, can simultaneously identify multiple types of objects, see Figure 6 shown.

[0032] S5, the diver or the submersible performs path navigation according to the diver's movement trajectory pushed in real time.

[0033] The present invention constructs a web interface to simulate a 3D scene, in which terrain models, display positions, movement trajectories, postures and other data can be loaded, and data is received and forwarded through the interaction between the front end and the back end. The WebSocket (a network communication protocol) is used to receive positioning data. It is a protocol for full-duplex communication on a single TCP connection, enabling real-time communication between the browser and the server without the need to send multiple HTTP requests to obtain data. This means that the connection is persistent and remains open until one party actively closes it.

[0034] Acoustic scanning is used to provide depth information (point cloud data) to assist in constructing the geometry of the three-dimensional space. By filtering, denoising, and feature extraction of the point cloud data, a clearer model can be generated. Algorithms (such as Poisson reconstruction, Delaunay triangulation, etc.) are used to transform the processed point cloud into a mesh model to generate a surface. The key to real-time modeling lies in being able to quickly respond to environmental changes. When the environment changes, the system only updates the affected parts, avoiding global reconstruction and improving efficiency. Data from multiple sensors is integrated to ensure the accuracy and integrity of the model. The generated three-dimensional model needs to have good visualization effects for easy user understanding and interaction. Real-time rendering technologies (such as shadows, lighting, materials, etc.) are applied to enhance the visual effects of the model. A real-time interaction function is provided, enabling users to rotate, zoom, and view the model from different angles.

[0035] Advanced computer vision technologies and deep learning models are used for environmental perception and intelligent target recognition to identify and locate various underwater objects, such as marine organisms, shipwrecks, lost items, etc. A classic deep learning object detection framework is adopted to quickly identify and locate multi-class objects in the input image. The system contains a specific underwater object dataset for training the model. This dataset may include annotated underwater pictures to help the model learn to distinguish different types of underwater targets and adapt to the unique imaging conditions of the water body, such as light refraction and color distortion.

[0036] The present invention realizes functions such as underwater operation scene restoration, real-time path navigation in the 3D scene, display of movement trajectories and postures, multi-view switching, real-time seabed modeling, rainbow diagram, environmental perception, and intelligent target recognition, making underwater operations safer and more efficient. At the same time, in complex environments and bad weather, an efficient and accurate underwater navigation solution is provided, which helps to improve operation efficiency and reduce errors.

[0037] The present invention has been described exemplarily in conjunction with the accompanying drawings. Obviously, the specific implementation of the present invention is not limited by the above-mentioned manner. As long as various non-substantive improvements are made by adopting the method concept and technical solution of the present invention, or the concept and technical solution of the present invention are directly applied to other occasions without improvement, they are all within the protection scope of the present invention.

Claims

1. A virtual reality navigation method, the specific steps of which are as follows: S1, build 3D scene; S2, loading the terrain model, as well as the data of the display position, movement trajectory, and posture of the diver or submersible in the 3D scene; S3, the terrain model is incrementally updated based on the real-time data from the acoustic scan; S4, the displayed position, movement trajectory, and posture of the diver or submersible are updated according to the real-time data pushed by the diver or submersible; S5: The diving commander makes judgments based on the diver’s or submersible’s position, movement trajectory, posture and other data to guide the operation.

2. A virtual reality navigation method according to claim 1, characterized in that: The three-dimensional scene construction in step S1 uses the JavaScript framework Vue.js to build the system front-end architecture and uses the JavaScript framework Three.js to build the 3D scene.

3. A virtual reality navigation method according to claim 1, characterized in that: The steps of building the terrain model in step S2 are as follows: S21, obtaining environmental information data through multi-beam scanning and refined scanning; S22, generating a three-dimensional point cloud through the environmental information data in step S21, and processing it using a library; S23, apply filtering algorithm to reduce the point cloud density, and use the feature calculation module provided by PCL to extract point cloud features; S24, using an algorithm to convert point cloud features into a three-dimensional mesh model; S25, using PCL or Open3D, applies color information to the generated mesh and uses UV mapping technology to attach the image to the 3D model surface.

4. A virtual reality navigation method according to claim 1, characterized in that: A rainbow map mode has been added to the terrain model. The specific implementation steps are as follows: S201, to obtain terrain elevation data; S202, calculating the overall height of the terrain, and setting a color dividing point using the ratio of the current height to the overall height; The formula is as follows: maxy = yArr.reduce((a, b) => (a > b ? a : b)); miny = yArr.reduce((a, b) => (a < b ? a : b)); height = maxy - miny; percent = (pos.getZ(i) - miny) / height; c1 = color1.clone().lerp(color2, percent * 1.5); c2 = color2.clone().lerp(color3, (percent - 0.32) * 1.7); Among them, maxy represents the maximum value of elevation data, a represents the current comparison data, b represents the next comparison data, miny represents the minimum value of elevation data, yArr represents the elevation data set, height represents the overall height of the mountain, percent represents the number of cycles of the ratio of the current height to the overall height, pos represents the terrain vertex data set, color1 represents the valley color, color2 represents the mountainside color, color3 represents the mountaintop color, c1 represents the interpolation between color1 and color2, c2 represents the interpolation between color2 and color3, reduce can iteratively compare the data of the elevation data set to find the maximum value, the getZ method can filter out the Z-axis data of the coordinate, the clone method is to copy the object, and lerp can regenerate the color value by linear interpolation; S203, setting the color attribute of the model data attributes in the form of color interpolation.

5. A virtual reality navigation method according to claim 4, characterized in that: The incremental update steps of the terrain model in step S3 are as follows: S31, performing deduplication check on the newly scanned real-time data; S32, merging the deduplicated real-time data into the existing terrain model; S33, incrementally updating the existing terrain model and switching it in the scene in real time.

6. A virtual reality navigation method according to claim 1, characterized in that: In step S2, the movement track of the diver or submersible is created using the built-in Line object of the JavaScript framework Three.js. First, a BufferGeometry and a LineBasicMaterial are created, and then a Line object is created to add the geometry and material to the object.

7. A virtual reality navigation method according to claim 6, characterized in that: The update formula of the action trajectory in step S4 is as follows: positions[i×3]=positionHistory[i].x positions[i×3+1]=positionHistory[i].y positions[i×3+2]=positionHistory[i].z maxPoints = positionHistory[i] + maxPoints Among them, positionHistory represents the array of stored points, maxPoints represents the maximum number, positions represents the vertex data, i represents the number of loop iterations, positionHistory[i].x represents getting the x-axis data of the current point, positionHistory[i].y represents getting the y-axis data of the current point, and positionHistory[i].z represents getting the z-axis data of the current point.

8. A virtual reality navigation method according to claim 1, characterized in that: The posture of the diver or submersible in step S2 is simulated by rotating the model position. The center point of the model needs to be calculated as the rotation center. The center point calculation formula is as follows: boundingBox=Box3(min,max) centerPosition=(boundingBox.min+boundingBox.max / 2,) Among them, boundingBox is the bounding box of the model calculated by THREE.Box3().setFromObject(data), min represents the minimum coordinate value, max represents the maximum value, BOX3 is the built-in object of the three.js model bounding box, and centerPosition represents the center point.

9. A virtual reality navigation method according to claim 1, characterized in that: The real-time data pushed by the diver or submersible in step S4 is obtained by collecting relevant data through the wave and water flow sensors, tide sensors and light sensors, and integrating the data from different sensors using data fusion technology to obtain the real-time status and real-time depth information of the current navigation environment.

10. A virtual reality navigation method according to claim 9, characterized in that: The real-time data pushed by the diver or submersible in step S4 also includes images. The images can be intelligently identified and located by the deep learning object detection framework YOLO. The specific steps are as follows: S41, YOLO treats the object detection task as a regression problem, directly predicting bounding boxes and class probabilities from image pixels; S42, divides the input image into S×S grids, each grid is responsible for predicting a fixed number of bounding boxes and their confidence and category probabilities; S43, the bounding box predicted by each grid consists of four coordinates and a confidence value, where the confidence value indicates the probability that the box contains an object; S44, uses the Softmax function to predict the category of each grid and can identify multiple categories of objects at the same time.