A Deep Learning-Based Automatic Orientation Method and Device for Containment Appearance Acquisition
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-30
- Publication Date
- 2026-08-14
AI Technical Summary
这个提取棱镜中心的方法受外部环境影响大,且精度一般
[0022]本发明的一种基于深度学习的安全壳外观采集设备的自动定向方案,实现了安全壳外观采集设备的零方向的精确且高效的解算。
Smart Images

Figure CN118379472B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of nuclear power plant containment inspection, and particularly relates to an automatic orientation technology solution for a containment appearance acquisition device based on deep learning. Background Technology
[0002] Due to environmental influences, nuclear power plant containment vessels may develop external defects after prolonged operation. To prevent further damage to the internal structure of the containment vessel, it is necessary to promptly identify these external defects and take effective reinforcement measures.
[0003] In the past, the visual inspection of containment structures primarily relied on manual inspection, a method that was time-consuming, risky, and inefficient. With technological advancements, close-range image acquisition devices have been applied to containment structure inspection. However, these devices are complex and heavy, making installation and transportation inconvenient. To address these issues, Wuhan University developed a remote acquisition device for containment structures. This device achieves visual inspection of containment structures through multi-site image acquisition and image stitching. The remote acquisition device boasts advantages such as light weight, simple structure, easy installation, and fast acquisition speed. It uses a pan-tilt unit (PTZ) to rotate the camera for shooting, but unlike a total station, the PTN lacks a directional telescope and automatic orientation function. To unify the PTN coordinate system across multiple stations and standardize image data from different stations to a single coordinate system during correction, the camera needs to be mounted at a known point, with that point used as the zero-direction orientation. Based on the direction of the line connecting the known mounting point and the orientation point, the rotation and translation relationships of the PTN coordinate system at different stations can be calculated, thus achieving a unified transformation. The method for automatic zero-direction orientation calculation of photographic equipment involves setting up the photographic equipment at a known point and a prism at another known point. The camera acquires the image of the prism, and the feature extraction method in image processing is used to automatically extract the center point of the prism, obtain the image-side coordinates of the center point, calculate the image-side offset distance between the center point and the principal point, calculate the object-side offset distance and angle between the prism center and the principal optical axis according to the principles of photogrammetry, and control the rotation of the pan-tilt head to make the prism center coincide with the principal optical axis, thus completing the zero-direction orientation of the photographic equipment.
[0004] In summary, a more accurate and efficient automatic orientation method is urgently needed for remote acquisition equipment within a containment structure. Traditional methods employ threshold segmentation, extracting three triangles based on the yellow-green color and area characteristics of the triangles on the prism target. A horizontal line is formed by connecting the vertices of the left and right triangles, and a vertical line is formed by passing through the vertex of the top triangle and perpendicular to the horizontal line. The intersection of the horizontal and vertical lines is the prism center. This method for extracting the prism center is highly susceptible to external environmental influences and has generally low accuracy. Summary of the Invention
[0005] The purpose of this invention is to solve the problem of automatic orientation of containment appearance acquisition equipment.
[0006] To overcome the shortcomings of the prior art, the technical solution proposed in this invention includes an automatic orientation method for a containment appearance acquisition device based on deep learning, comprising the following steps.
[0007] Step 1: Set up the data acquisition equipment and directional prism at the two known points respectively;
[0008] Step 2: Control the containment appearance acquisition device to roughly aim at the directional prism and take a picture of the directional prism;
[0009] Step 3: Pre-collect images of directional prisms under different environments to construct a prism recognition dataset. Crop out images of prism regions from the dataset to construct a center recognition dataset. Construct an automatic prism center detection model. The prism center detection model consists of two stages. In the first stage, identify prisms in the images and crop out the detected local images containing prisms to proceed to the second stage, where the center of the prism is further identified from the local images. Train the automatic prism center detection model using the prism recognition dataset and the center recognition dataset.
[0010] Step 4: Input the image of the directional prism taken in Step 2 into the prism center automatic detection model trained in Step 3, extract the image-side coordinates of the prism center point, and calculate its image-side offset distance from the principal image point.
[0011] Step 5: Calculate the object-side offset distance and offset angle between the center of the orientation marker and the principal image point. Control the pan-tilt unit to rotate so that the center of the orientation marker coincides with the principal image point to accurately aim at the prism. After re-aiming, take another picture of the prism to check whether the center of the prism coincides with the center of the principal image point.
[0012] Furthermore, image enhancement was used to expand the prism recognition dataset and the center recognition dataset respectively.
[0013] Moreover, the first stage uses the traditional YOLOv8 network model.
[0014] Moreover, the second stage adopts an improved YOLOv8 network model. The improved YOLOv8 network model is based on the traditional YOLOv8 network structure, with the addition of a multi-scale fusion framework to improve the model's feature extraction capability. At the same time, it also adds a multi-feature encoding module and a small object detection head to improve the detection performance of small targets such as the center of a prism.
[0015] Furthermore, the improved YOLOv8 model is used to detect the center of the prism, resulting in a rectangular frame. The center of the rectangular frame is used as the center of the prism, and the corresponding pixel coordinates are returned. Then, compared with the principal point coordinates, the horizontal and vertical offset distances and offset angles of the object are calculated according to the principles of photogrammetry. Based on the calculated offset angles, the gimbal is controlled to re-aimine and complete the orientation of the acquisition device.
[0016] Furthermore, the enclosure appearance acquisition device includes a tripod, a base, a gimbal mounting base, a high-precision gimbal, a camera connection plate, and a camera. The camera is connected to the high-precision gimbal via the camera connection plate. The gimbal mounting base is then connected to the bottom of the high-precision gimbal, and the gimbal mounting base is connected to the base. Finally, the entire device is fixed to the tripod via the base.
[0017] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the automatic orientation method of the deep learning-based safety shell appearance acquisition device as described above.
[0018] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the automatic orientation method for a deep learning-based containment appearance acquisition device as described above.
[0019] On the other hand, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the automatic orientation method for a deep learning-based security shell appearance acquisition device as described above.
[0020] Furthermore, it includes a processor and a memory, the memory being used to store program instructions, and the processor being used to call the stored instructions in the memory to execute the automatic orientation method of a deep learning-based containment appearance acquisition device as described above.
[0021] Alternatively, it may include a readable storage medium storing a computer program that, when executed, implements an automatic orientation method for a deep learning-based enclosure appearance acquisition device as described above.
[0022] The present invention provides an automatic orientation scheme for a containment appearance acquisition device based on deep learning, which realizes accurate and efficient zero-direction calculation of the containment appearance acquisition device.
[0023] The present invention is simple and convenient to implement, highly practical, and solves the problems of low practicality and inconvenience in actual application of related technologies. It can improve user experience and has significant market value. Attached Figure Description
[0024] Figure 1 This is an overall flowchart of the automatic orientation of the containment appearance acquisition device according to an embodiment of the present invention;
[0025] Figure 2 This is a structural diagram of a YOLOv8 prism for detection according to an embodiment of the present invention;
[0026] Figure 3 This is a structural diagram of the improved YOLOv8 for detecting the center of a prism according to an embodiment of the present invention;
[0027] Figure 4 This is a schematic diagram of the detection process according to an embodiment of the present invention.
[0028] Figure 5 This is a schematic diagram of the offset of the positioning result in an embodiment of the present invention. Detailed Implementation
[0029] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0030] This invention discloses an automatic orientation scheme for a safety shell appearance acquisition device based on deep learning. First, the acquisition device and an orientation prism are set up at two known points. Then, the pan-tilt unit is controlled to roughly align the camera with the orientation prism, capturing images of the prism. The acquired images are used to construct prism recognition datasets and center recognition datasets. After the datasets are created, they are used to train a YOLOv8 model and an improved YOLOv8 model, respectively. The trained YOLOv8 model and the improved YOLOv8 model are then used for the orientation of the acquisition device. After the acquisition device roughly aligns with the prism and captures images, the images are transmitted to a computer. The computer uses the trained models to first identify the prism in the image, and then detects the prism center in the identified prism image, using the center of the detection box as the pixel coordinates of the detected prism center. The image-side offset distance between the prism and the principal point is calculated, and the object-side offset distance and angle between the prism center and the principal optical axis are calculated according to photogrammetry principles. The pan-tilt unit is then rotated to align the prism center with the principal optical axis, completing the zero-direction orientation of the photographic device.
[0031] This invention proposes an overall process for the automatic orientation of a containment structure appearance acquisition device based on deep learning, as follows: Figure 1 As shown, the specific steps are as follows:
[0032] S1. Set up the data acquisition equipment and orientation prism at the known points;
[0033] This step involves setting up the data acquisition equipment and the directional prism at two known points, as implemented in the example below:
[0034] (1) Set up two markers on the ground as station points and use a total station to obtain the coordinate information of the two points.
[0035] (2) Set up the acquisition equipment and orientation prism on the known point using a tripod, and complete the centering and leveling operation.
[0036] In practice, a total station can be used to obtain the coordinates of two points first. Then, the acquisition equipment and prism are set up at the two points respectively. The acquisition equipment includes a tripod, base, pan-tilt head mounting base, high-precision pan-tilt head, camera connection plate, and camera. The camera is connected to the high-precision pan-tilt head via the camera connection plate, and then the pan-tilt head mounting base is connected to the bottom of the high-precision pan-tilt head. The pan-tilt head mounting base is connected to the base, and finally, the entire assembly is fixed to the tripod via the base.
[0037] S2. Control the acquisition device to capture images of the prism;
[0038] This step involves controlling the acquisition device to rotate and roughly align it with the directional prism, then taking an image of the directional prism. In practice, a computer can be connected to a high-precision pan-tilt unit and camera via a wired connection. The pan-tilt unit is then rotated to align the camera with the prism, and the camera takes a picture, which is then transmitted back to the computer in real time.
[0039] The implementation in the example is as follows:
[0040] (1) Connect the gimbal and camera to the laptop via a data transmission cable.
[0041] (2) Control the rotation of the gimbal so that the prism enters the shooting range of the camera.
[0042] (3) Control the camera to capture prism images and send them back to the laptop.
[0043] S3. Construct the dataset and network model;
[0044] This step involves pre-collecting images of directional prisms under different environments to establish a prism center recognition dataset and construct an automatic prism center detection model. The prism center detection model comprises two stages: the first stage identifies the prism from a large image, obtaining a local image containing the prism; the second stage then accurately identifies the prism center from this local image. The purpose of the first stage is to narrow down the detection image area, improving the detection accuracy and efficiency of the second stage.
[0045] The first stage preferably uses the YOLOv8 network model to identify prisms in the image, and then crops out the local image containing the prisms to proceed to the second stage. The second stage preferably uses an improved YOLOv8 network model to further identify the center of the prism from the local image.
[0046] The prism center recognition dataset includes a prism recognition dataset and a center recognition dataset. The prism recognition dataset is used to train YOLOv8 to identify prisms from the whole image. The center recognition dataset is used to train an improved YOLOv8 model to further detect the center of the prism from the identified prism images. The YOLOv8 model described in the first stage of this invention is the traditional YOLOv8 network structure, an improvement on YOLOv5. This YOLOv8 model uses the more efficient C2f module to replace the C3 module and adjusts the number of channels. The detection head is also modified to use decoupled head technology for classification and detection. The YOLOv8 network structure is simplified, resulting in faster detection speed and higher detection accuracy. Meanwhile, the improved YOLOv8 model for detecting prism centers incorporates a multi-scale fusion framework on top of the traditional YOLOv8 network structure, improving the model's feature extraction capabilities. It also adds a multi-feature encoding module and a small object detection head to improve the detection performance for small targets such as prism centers.
[0047] Furthermore, this invention proposes that after acquiring and processing the prism image dataset, image enhancement techniques be used to augment the dataset images.
[0048] The preferred implementation method used in the embodiments is as follows:
[0049] (1) The original images collected in this embodiment have a resolution of 6000×4000. The prism center occupies a small proportion of the original image, and directly identifying the center on the original image results in poor accuracy. Therefore, this invention uses a two-stage detection model to identify the prism center. To improve the usability of the model, this embodiment of the invention collected prism images in different environments and constructed a prism recognition dataset with an image resolution of 6000×4000. Then, images of the prism region were cropped from the images to construct a center recognition dataset with a resolution of 320×320. Data augmentation methods were used, preferably by adding image noise and changing image brightness to simulate different weather conditions, to expand the two datasets and improve the training effect of the model.
[0050] (2) Construct a prism recognition model. It is recommended to use the YOLOv8 model from existing technologies for the preferred detection model. For ease of implementation and reference, the model structure is provided as follows: Figure 2As shown, the model consists of four parts: input, skeleton network, neck network, and detection head. The input image size in the prism recognition model is 6000×4000. The model includes a total of 23 network layers. Layer 1: A CBS module is used to convert the input feature map to 64 channels with a stride of 2, resulting in the output feature map P1. Layer 2: The CBS module is used again to convert the input feature map from 64 channels to 128 channels with a stride of 2, resulting in the output feature map P2. Layer 3: The C2f module is used, and the output result has 128 channels. Layer 4: The CBS module is used again to convert the input feature map from 128 channels to 256 channels with a stride of 2, resulting in the output feature map P3. Layer 5: The C2f module is used, and the output result has 256 channels. Layer 6: The CBS module is used again to convert the input feature map from 256 channels to 512 channels with a stride of 2, resulting in the output feature map P4. Layer 7: Using the C2f module, the number of channels in the input feature map is adjusted to 512. Layer 8: Using the CBS module, the input feature map is converted from 512 channels to 1024 channels. Layer 9: Using the C2f module, the number of channels in the input feature map is adjusted to 1024. Then, the SPPF module in layer 10 is used to process the feature map.
[0051] Layer 11: Upsample the feature map using an upsample operation, doubling its size. Layer 12: Concatenate the feature map from the previous step with the feature map generated in Layer 7. Layer 13: Process the feature map using the C2f module, adjusting the number of channels in the input feature map to 512. Layer 14: Upsample the feature map using an upsample operation, doubling its size. Layer 15: Concatenate the feature map from the previous step with the feature map generated in Layer 5. Layer 16: Process the feature map using the C2f module, adjusting the number of channels in the input feature map to 256, resulting in a small-sized feature map. Layer 17: Use the CBS module, outputting a 256-channel feature map. Layer 18: Concatenate the feature map from the previous step with the feature map generated in Layer 13. Layer 19: Process the feature map using the C2f module, adjusting the number of channels in the input feature map to 512, resulting in a medium-sized feature map. Layer 20: Use the CBS module, outputting a 512-channel feature map. Layer 21: The feature map from the previous step is concatenated with the feature map generated in layer 10. Layer 22: The C2f module is used to process the feature map, adjusting the number of channels in the input feature map to 512 to obtain a large-scale feature map. Layer 23: Finally, the feature maps at three scales are input into the Detect module for object detection. YOLOv8 also replaces the original coupling head with the currently mainstream decoupling head, separating the classification head and the regression detection head.
[0052] The structure of the CBS module is as follows: Figure 2 As shown, it consists of convolutional layers (Conv2d), normalization layers (BatchNorm), and SiLU activation functions. The structure of the C2f module is as follows: Figure 2 As shown, this module implements cross-stage feature fusion and partial feature reuse. Through the organic combination of components such as convolutional layers (Conv2d), bottleneck blocks (Bottleneck), skip connections (split), and feature fusion (concat), it effectively integrates low-level and high-level features, improving the network's perceptual and generalization capabilities and providing richer and more useful feature representations for object detection tasks. Each bottleneck block in the C2f module typically consists of a series of convolutional layers, including 1x1 convolutions, 3x3 convolutions, etc. A bottleneck structure is usually adopted, i.e., first reducing the number of channels, then performing feature extraction through 3x3 convolutions, and finally restoring the number of channels. This helps reduce computational cost while enhancing feature representation capabilities. The Concat module is a feature fusion operation that concatenates two tensors along a specified dimension to generate a new tensor. This operation is typically used to fuse feature maps with different numbers of channels to expand the representational power of features. The SPPF module divides the feature map into multiple grids of different scales, performs max pooling on each grid, and finally concatenates all pooling results through a concat layer to form a fixed-length feature vector, thus providing more comprehensive spatial information.
[0053] (3) Constructing a center recognition model: The detection model used in this invention is based on the YOLOv8 model with corresponding improvements to better achieve center recognition. Specifically, by fusing feature maps of different scales, local and global feature information is better obtained. A Zoom_cat module is constructed to spatially stitch features of different sizes to capture the local fine details of small target objects, improving the detection effect for small targets such as prism centers. The model structure is as follows: Figure 3As shown. The improved YOLOv8 has 31 layers, with 10 layers in the backbone network identical to YOLOv8. Layers 11 and 12 are CBS modules, performing 1x1 convolution operations on the features. Layer 11 also processes features from layer 5 of the backbone network. Layer 13 is a Zoom_cat module, concatenating features from layers 11, 12, and 7 at the channel level. Layer 14 is a C2f module, extracting and transforming features from the concatenated features. Layers 15 and 16 are CBS modules. The CBS module in layer 16 processes features from layer 3 of the backbone network, while the Zoom_cat module in layer 17 concatenates features from layers 16, 5, and 15 at the channel level. Layer 18 is a C2f module, extracting and transforming features from the concatenated features. Layer 19's CBS module processes features from layer 18. Layer 20 is a Concat module, fusing features from layers 19 and 15. Layer 21 is the C2f module, which processes the features from layer 20. Layer 22 is the CBS module, which performs a 3x3 convolution operation on the features from layer 21 using a convolutional layer. Then, layer 23's Concat module concatenates the features from this layer with the features from layer 15 in the backbone network at the channel level, increasing the number of channels. Next, layer 24 uses the C2f module to extract and transform the concatenated features, keeping the number of channels at 512 and the feature map size unchanged. Then, layer 25 uses the ScalSeq module to scale the features from layers 5, 7, and 9 in the backbone network, adjusting the number of channels to 256. Then, layer 26 uses the Add module to add the features adjusted by the ScalSeq module to the features output from layer 18 element-wise to enhance feature representation. Next, layer 27 performs an upsampling operation using Upsample, doubling the feature map size. Finally, layer 28 uses the Concat module to concatenate the upsampled features with the features from layer 3 in the backbone network at the channel level, increasing the number of channels. Then, layer 29 uses the C2f module to extract and transform the concatenated features, adjusting the number of channels to 128 while keeping the feature map size unchanged. Finally, layer 30 uses the ScalSeq module to scale the features from layers 3, 26, and 21, adjusting the number of channels to 128. In layer 31, the Add module is used to element-wise add the features adjusted by the ScalSeq module to the features output by the C2f module in layer 29, enhancing feature representation. For the detection head (Detect), the original YOLOv8 structure only has three detection heads. To improve the detection performance for small targets, a new detection head specifically for small targets is added. Similar to YOLOv8, each detection head also uses a decoupled head consisting of two convolutional modules to extract the target location and category information separately.The detection head for small targets, along with the other three detection heads, all use the same decoupling head structure, but the input data is different.
[0054] The Add operation adds two tensors element-wise to generate a new tensor. This operation is typically used to fuse feature maps with the same spatial dimensions to enhance the correlation and complementarity between features. The ScalSeq structure is as follows: Figure 3 As shown, the function first receives three input feature maps, representing features from different levels. Then, it uses three 1x1 convolutional layers (Conv2d) to adjust the channel count of two of the features. Next, it uses interpolation to match the adjusted feature size with the third feature. Then, it concatenates the three features along the channel dimension (Concat), further fuses them using a 3D convolutional layer (Conv3d), then performs feature enhancement using batch normalization (BatchNorm) and an activation function (Silu), and finally downsamples the features using 3D max pooling (Pool3d) to generate the final output features. Layers 13 and 17 preferably use the Zoom_cat module to fuse feature maps from different levels. The specific structure of the Zoom_cat module is existing technology and will not be described in detail here. The input is a list of three feature maps, representing feature maps from large, medium, and small sizes, respectively. Then, during the forward propagation, the large feature map is first adjusted to the same size as the medium feature map using adaptive pooling. Next, the adjusted large feature map is added to the results of its max pooling and average pooling to enhance the features. Then, interpolation is used to adjust the small feature map to the same size as the medium feature map. Finally, the adjusted large, medium, and small feature maps are concatenated along the channel dimension, and the concatenated feature map is returned.
[0055] (4) Training the model: After constructing the prism recognition model and the center recognition model, the two models were trained using the prism recognition dataset and the center recognition dataset, respectively.
[0056] S4. Obtain the image-side coordinates of the prism center;
[0057] The testing process is as follows Figure 4 As shown, after obtaining the two trained models, the image captured by the acquisition device in step 2, initially aimed at the prism, can be input into the prism recognition model to detect the prism. The prism image within the detection box is then extracted from the original image and transmitted to the center recognition model. The result obtained by the center detection model is a small rectangular box that selects the center of the prism. Using the center of this rectangle as the center of the prism, the image-side coordinates of the prism center on the original image are obtained.
[0058] S5. Calculate the offset angle;
[0059] Because the improved YOLOv8 model is used to detect the prism center, the result is a rectangular bounding box. The center of this bounding box is used as the prism center, and the corresponding pixel coordinates (u, v) are returned. Then, these coordinates are compared with the principal point coordinates (u, v). o ,v o In contrast, based on the principles of photogrammetry, the horizontal and vertical offset distances L of the object can be calculated. x ,L y and offset angle θ x ,θ y .
[0060] In this embodiment, after obtaining the image-side coordinates, the offset between that point and the principal image point is calculated. The pixel offsets in the horizontal and vertical directions are obtained separately. Based on photogrammetry principles, the horizontal and vertical offset distances L in the object-side direction are calculated. x ,L y and offset angle θ x ,θ y The calculation formula is as follows:
[0061]
[0062]
[0063]
[0064]
[0065] Where (u,v) are pixel coordinates, (u o ,v o ) represents the coordinates of the principal point, L x ,L y θ represents the horizontal and vertical offset distances of the object, respectively. x ,θ y Let D be the offset angle, L be the shooting distance, f be the principal distance, w be the camera sensor width, and W be the image frame width. The orientation of the acquisition device can be completed by controlling the pan-tilt unit to re-aim based on the calculated offset angle.
[0066] In practice, an automatic prism center detection model can be trained in advance using prism recognition datasets and center recognition datasets. When actual orientation is required, the orientation prism is roughly aimed at using the enclosure appearance acquisition device, and a photo is taken. The photo is then input into the trained automatic prism center detection model, which uses a deep learning-based detection method to obtain the center of the prism in the image and calculates the horizontal and vertical offset angles. Based on the calculation results, the gimbal is rotated to accurately aim at the prism. After re-aiming, the prism is photographed again to check whether the prism center coincides with the center of the principal point of the image.
[0067] In specific implementation, the method proposed in the technical solution of this invention can be automatically executed by those skilled in the art using computer software technology. System devices for implementing the method, such as computer-readable storage media storing the corresponding computer program of the technical solution of this invention and computer equipment including the computer program running the corresponding computer program, should also be within the protection scope of this invention.
[0068] The automatic orientation device for the containment appearance acquisition device based on deep learning provided by the present invention is described below. The automatic orientation device for the containment appearance acquisition device based on deep learning described below and the automatic orientation method for the containment appearance acquisition device based on deep learning described above can be referred to in correspondence with each other.
[0069] The electronic device provided by this invention may include: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus. The processor can invoke logical instructions from the memory to execute an automatic orientation method for a deep learning-based containment appearance acquisition device, the method comprising:
[0070] Step 1: Set up the data acquisition equipment and directional prism at the two known points respectively;
[0071] Step 2: Control the containment appearance acquisition device to roughly aim at the directional prism and take a picture of the directional prism;
[0072] Step 3: Pre-collect images of directional prisms under different environments to construct a prism recognition dataset. Crop out images of prism regions from the dataset to construct a center recognition dataset. Construct an automatic prism center detection model. The prism center detection model consists of two stages. In the first stage, identify prisms in the images and crop out the detected local images containing prisms to proceed to the second stage, where the center of the prism is further identified from the local images. Train the automatic prism center detection model using the prism recognition dataset and the center recognition dataset.
[0073] Step 4: Input the image of the directional prism taken in Step 2 into the prism center automatic detection model trained in Step 3, extract the image-side coordinates of the prism center point, and calculate its image-side offset distance from the principal image point.
[0074] Step 5: Calculate the object-side offset distance and offset angle between the center of the orientation marker and the principal image point. Control the pan-tilt unit to rotate so that the center of the orientation marker coincides with the principal image point to accurately aim at the prism. After re-aiming, take another picture of the prism to check whether the center of the prism coincides with the center of the principal image point.
[0075] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0076] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium, wherein when the computer program is executed by a processor, the computer is able to execute the automatic orientation method for a security shell appearance acquisition device based on deep learning provided by the above methods, the method comprising:
[0077] Step 1: Set up the data acquisition equipment and directional prism at the two known points respectively;
[0078] Step 2: Control the containment appearance acquisition device to roughly aim at the directional prism and take a picture of the directional prism;
[0079] Step 3: Pre-collect images of directional prisms under different environments to construct a prism recognition dataset. Crop out images of prism regions from the dataset to construct a center recognition dataset. Construct an automatic prism center detection model. The prism center detection model consists of two stages. In the first stage, identify prisms in the images and crop out the detected local images containing prisms to proceed to the second stage, where the center of the prism is further identified from the local images. Train the automatic prism center detection model using the prism recognition dataset and the center recognition dataset.
[0080] Step 4: Input the image of the directional prism taken in Step 2 into the prism center automatic detection model trained in Step 3, extract the image-side coordinates of the prism center point, and calculate its image-side offset distance from the principal image point.
[0081] Step 5: Calculate the object-side offset distance and offset angle between the center of the orientation marker and the principal image point. Control the pan-tilt unit to rotate so that the center of the orientation marker coincides with the principal image point to accurately aim at the prism. After re-aiming, take another picture of the prism to check whether the center of the prism coincides with the center of the principal image point.
[0082] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements an automatic orientation method for a deep learning-based containment appearance acquisition device provided by the methods described above, the method comprising:
[0083] Step 1: Set up the data acquisition equipment and directional prism at the two known points respectively;
[0084] Step 2: Control the containment appearance acquisition device to roughly aim at the directional prism and take a picture of the directional prism;
[0085] Step 3: Pre-collect images of directional prisms under different environments to construct a prism recognition dataset. Crop out images of prism regions from the dataset to construct a center recognition dataset. Construct an automatic prism center detection model. The prism center detection model consists of two stages. In the first stage, identify prisms in the images and crop out the detected local images containing prisms to proceed to the second stage, where the center of the prism is further identified from the local images. Train the automatic prism center detection model using the prism recognition dataset and the center recognition dataset.
[0086] Step 4: Input the image of the directional prism taken in Step 2 into the prism center automatic detection model trained in Step 3, extract the image-side coordinates of the prism center point, and calculate its image-side offset distance from the principal image point.
[0087] Step 5: Calculate the object-side offset distance and offset angle between the center of the orientation marker and the principal image point. Control the pan-tilt unit to rotate so that the center of the orientation marker coincides with the principal image point to accurately aim at the prism. After re-aiming, take another picture of the prism to check whether the center of the prism coincides with the center of the principal image point.
[0088] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0089] To verify the superior performance of the proposed model on the prism center point detection task, YOLOv8s, YOLOv7, YOLOv5s, YOLOv3, and Centernet models were trained and tested on the prism center point dataset, and their test results were compared with those of the proposed model. To ensure a fair comparison, the operating environment parameters were kept consistent during training, and all methods were trained until they converged to achieve optimal performance. Table 1 shows the detection results of each model on the test set. As can be seen from Table 1, except for the proposed model, the YOLOv7 model performed best, with an F1 score of 90.66%, which is the best among all models, but its parameters and GFLOPs are relatively large. The proposed model outperforms these models in terms of mAP@0.5 and mAP@0.5:0.95. Compared with YOLOv7, mAP@0.5 and mAP_0.5-0.95 are improved by 1.68% and 2.01%, respectively, while the parameters and GFLOPs are also significantly reduced. These results also demonstrate that the network model proposed in this invention has good detection performance in the prism center detection task.
[0090] Table 1. Comparison of detection results from different models.
[0091]
[0092] To verify the accuracy of the prism center detected by the proposed method, automatic detection results under different environments were compared with manual calibration results, and the offset pixel values between the two were calculated. The results are as follows: Figure 5 As shown. From Figure 5 As can be seen, the vast majority of offsets are within 2 pixels. Table 2 summarizes the error information of the test results. The maximum offsets are -2.688 pixels horizontally and -2.188 pixels vertically, respectively, and the root mean square error values are 1.079 pixels horizontally and 0.923 pixels vertically, respectively. The results show that the method proposed in this invention can accurately locate the pixel coordinates of the prism center point. Converted to angles using the angle calculation formula, the detection deviations for horizontal and vertical angles are 4.51″ and 3.67″, respectively, both less than the minimum rotation angle of the high-precision gimbal, and the accuracy is far higher than the requirements of practical applications.
[0093] Table 2. Positioning Error Information
[0094]
[0095] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
[0097] The above embodiments are merely illustrative examples of the technical solutions of the present invention. The mobile edge computing task scheduling and offloading method, apparatus, and storage medium based on multi-agent collaborative deep reinforcement learning involved in the present invention are not limited to the contents described in the above embodiments, but are subject to the scope defined by the claims. Any modifications, additions, or equivalent substitutions made by those skilled in the art based on these embodiments are within the scope of protection claimed by the claims of the present invention.
Claims
1. An automatic orientation method for a containment appearance acquisition device based on deep learning, characterized in that: Includes the following steps, Step 1: Set up the data acquisition equipment and directional prism at the two known points respectively; Step 2: Control the containment appearance acquisition device to roughly aim at the directional prism and take a picture of the directional prism; Step 3: Pre-collect images of directional prisms under different environments to construct a prism recognition dataset. Crop out images of prism regions from the dataset to construct a center recognition dataset. Construct an automatic prism center detection model. The prism center detection model consists of two stages. In the first stage, identify prisms in the images and crop out the detected local images containing prisms to proceed to the second stage, where the center of the prism is further identified from the local images. Train the automatic prism center detection model using the prism recognition dataset and the center recognition dataset. Step 4: Input the image of the directional prism taken in Step 2 into the prism center automatic detection model trained in Step 3, extract the image-side coordinates of the prism center point, and calculate its image-side offset distance from the principal image point. Step 5: Calculate the object-space offset distance and offset angle between the orientation marker center and the principal image point. Control the pan-tilt unit to rotate so that the orientation marker center coincides with the principal image point to accurately aim at the prism. After re-aiming, take another picture of the prism to check whether the prism center coincides with the principal image point center. This process includes the following steps. After obtaining the image coordinates, the offset from the principal point is calculated, yielding the pixel offsets in both the horizontal and vertical directions. Based on photogrammetry principles, the horizontal and vertical offset distances in the object plane are then calculated. , and offset angle , as follows, in,( , ) represents pixel coordinates, ( , ( ) represents the coordinates of the principal point. , These are the horizontal and vertical offset distances of the object, respectively. , For offset angle, For photography distance, This is the offset distance. Main distance, For the camera sensor width, Image width; The gimbal is controlled to re-aim at the calculated offset angle to complete the orientation of the acquisition device.
2. The automatic orientation method for the safety shell appearance acquisition device based on deep learning according to claim 1, characterized in that: Image enhancement was used to expand the prism recognition dataset and the center recognition dataset respectively.
3. The automatic orientation method for the safety shell appearance acquisition device based on deep learning according to claim 1, characterized in that: The first stage uses the traditional YOLOv8 network model.
4. The automatic orientation method for a deep learning-based containment appearance acquisition device according to claim 1, characterized in that: The second stage employs an improved YOLOv8 network model. This improved YOLOv8 network model is based on the traditional YOLOv8 network structure, with the addition of a multi-scale fusion framework to enhance the model's feature extraction capabilities. It also adds a multi-feature encoding module and a small object detection head to improve the detection performance of small targets such as the center of a prism.
5. The automatic orientation method for the safety case appearance acquisition device based on deep learning according to claim 1, characterized in that: The improved YOLOv8 model is used to detect the center of the prism, and the result is a rectangular frame. The center of the rectangular frame is used as the center of the prism, and the corresponding pixel coordinates are returned. Then, compared with the principal point coordinates, the horizontal and vertical offset distances and offset angles of the object are calculated according to the principles of photogrammetry. Based on the calculated offset angle, the gimbal is controlled to re-aim and complete the orientation of the acquisition device.
6. The automatic orientation method for a deep learning-based containment appearance acquisition device according to claim 1, 2, 3, 4, or 5, characterized in that: The enclosure appearance acquisition device includes a tripod, a base, a gimbal mounting base, a high-precision gimbal, a camera connection plate, and a camera. The camera is connected to the high-precision gimbal via the camera connection plate. The gimbal mounting base is then connected to the bottom of the high-precision gimbal and connected to the base. Finally, the entire device is fixed to the tripod via the base.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it implements the automatic orientation method for the deep learning-based enclosure appearance acquisition device as described in any one of claims 1 to 6.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the automatic orientation method for the deep learning-based enclosure appearance acquisition device as described in any one of claims 1 to 6.
9. A computer program product, comprising a computer program, characterized in that: When the computer program is executed by the processor, it implements the automatic orientation method for the deep learning-based enclosure appearance acquisition device as described in any one of claims 1 to 6.