Panoramic image generation method and device, ship and medium
By combining lidar, dual-light pods, and millimeter-wave radar, multi-source, multi-modal data is acquired, solving the problem of insufficient accuracy and reliability in panoramic image generation in existing technologies, and realizing high-precision panoramic image generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NAT ENG RES CENT OF DREDGING TECH & EQUIP
- Filing Date
- 2026-04-07
- Publication Date
- 2026-05-05
AI Technical Summary
Existing panoramic image generation methods struggle to accurately determine the distance between ships and obstacles, as well as the three-dimensional shape of obstacles, resulting in insufficient accuracy and reliability of the generated panoramic images.
Point cloud data is acquired using LiDAR, dual-light pods acquire dual-light image data from multiple viewpoints, and millimeter-wave radar acquires obstacle detection data. Panoramic images are generated through depth extraction, clustering, and feature extraction, enabling multi-source and multi-modal perception, accurate alignment of image data and point cloud data, and fusion of obstacle detection results.
It improves the accuracy and reliability of panoramic images, comprehensively acquires environmental information, makes up for the shortcomings of a single sensor, and achieves high-precision and highly robust panoramic image generation.
Smart Images

Figure CN121982125A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of image technology, and in particular to a method, apparatus, ship, and medium for generating panoramic images. Background Technology
[0002] With the continuous improvement of automation and intelligence in ship operations, comprehensive and accurate perception of the operating environment has become crucial to ensuring navigation and operational safety. Especially in complex waters, port operations, and offshore engineering projects, ships need to have real-time, all-round access to information about their surrounding environment. Therefore, generating panoramic images of the ship's operating area (i.e., the surrounding environment) has become an important means of achieving comprehensive environmental perception, and its necessity and importance are increasingly prominent.
[0003] Currently, panoramic image generation methods mainly employ image stitching techniques based on a single visual sensor (such as a monocular camera or a multi-view camera array). However, this method struggles to accurately determine the distance between the ship and obstacles, as well as the three-dimensional shape of the obstacles, resulting in insufficient accuracy and reliability of the generated panoramic images.
[0004] Therefore, there is an urgent need to propose a new method to solve the above problems. Summary of the Invention
[0005] This invention provides a method, apparatus, vessel, and medium for generating panoramic images, which can effectively improve the accuracy and reliability of panoramic image generation.
[0006] In a first aspect, embodiments of the present invention provide a method for generating panoramic images, applied to a ship, wherein a lidar and a dual-light pod are installed on the top of the ship, and millimeter-wave radars are installed on the bow and stern of the ship, respectively. The method includes:
[0007] The laser radar acquires point cloud data of the ship's operating area, the dual-light pod acquires dual-light image data of the ship's operating area from multiple perspectives, and the millimeter-wave radar acquires obstacle detection data of the ship's operating area.
[0008] The point cloud data is subjected to depth extraction processing to obtain pixel-level depth information, and the dual-light image data is subjected to coordinate transformation processing based on the pixel-level depth information to obtain target visual data.
[0009] Clustering is performed on the point cloud data to obtain obstacle point cloud clusters, and feature extraction is performed on the target visual data to obtain visual target detection results;
[0010] Based on the obstacle detection data, the obstacle point cloud clusters, and the visual target detection results, a current panoramic image of the ship's operating area is generated.
[0011] The technical solution of this invention first acquires point cloud data of the ship's operating area using LiDAR, then acquires dual-light image data of the operating area from multiple perspectives using a dual-light pod, and finally acquires obstacle detection data of the operating area using millimeter-wave radar. This allows for comprehensive acquisition of environmental information of the operating area from multiple dimensions, including three-dimensional spatial structure, multispectral visual images, and obstacle location information. This achieves multi-source, multi-modal, and omnidirectional perception of the operating area, effectively compensating for the shortcomings of single sensors in imaging, ranging, or obstacle detection, and improving the completeness and reliability of environmental perception. It provides a rich and accurate data foundation for the subsequent generation of high-precision, robust panoramic images. Next, depth extraction processing is performed on the point cloud data to obtain pixel-level depth information. Based on this pixel-level depth information, coordinate transformation processing is performed on the dual-light image data to obtain target visual data. This achieves precise spatial alignment between the image data and the point cloud data, providing a rigorous registration data foundation for the subsequent generation of high-precision panoramic images. Subsequently, the point cloud data is clustered to obtain obstacle point cloud clusters, and feature extraction is performed on the target visual data to obtain visual target detection results. This improves the accuracy and effectiveness of obstacle localization, clarifies the specific shape and visual attributes of visual targets, and achieves precise obstacle classification and image-level localization. It overcomes the limitations of single-modal visual detection and provides reliable data support for the subsequent generation of panoramic images. Finally, based on the obstacle detection data, obstacle point cloud clusters, and visual target detection results, a current panoramic image of the ship's operating area is generated. This achieves precise overlay and fusion of obstacle detection data, obstacle point cloud clusters, and visual target detection results in the panoramic image, significantly improving the accuracy, reliability, and information richness of the generated panoramic image. Therefore, the technical solution of this invention solves the problem in the prior art where it is difficult to accurately determine the distance between the ship and obstacles and the three-dimensional shape of the obstacles, resulting in insufficient accuracy and reliability of the generated panoramic image.
[0012] Secondly, embodiments of the present invention also provide a panoramic image generation device applied to a ship, wherein a lidar and a dual-light pod are installed on the top of the ship, and millimeter-wave radars are respectively installed on the bow and stern of the ship. The device includes:
[0013] The acquisition module is used to acquire point cloud data of the ship's operating area through the lidar, acquire dual-light image data of the ship's operating area from multiple perspectives through the dual-light pod, and acquire obstacle detection data of the ship's operating area through the millimeter-wave radar.
[0014] The conversion module is used to perform depth extraction processing on the point cloud data to obtain pixel-level depth information, and to perform coordinate transformation processing on the dual-light image data based on the pixel-level depth information to obtain target visual data.
[0015] The extraction module is used to perform clustering processing on the point cloud data to obtain obstacle point cloud clusters, and to extract features from the target visual data to obtain visual target detection results;
[0016] The generation module is used to generate a current panoramic image of the ship's operating area based on the obstacle detection data, the obstacle point cloud clusters, and the visual target detection results.
[0017] Thirdly, embodiments of the present invention also provide a vessel, the vessel comprising:
[0018] At least one processor; and a memory communicatively connected to said at least one processor;
[0019] The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the panoramic image generation method according to any embodiment of the present invention.
[0020] Fourthly, embodiments of the present invention also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, implement the panoramic image generation method described in any embodiment of the present invention.
[0021] Fifthly, the present invention provides a computer program product comprising computer instructions that, when executed on a computer, cause the computer to perform the method for generating panoramic images as provided in the first aspect.
[0022] It should be noted that the aforementioned computer instructions may be stored, in whole or in part, on a computer-readable storage medium. This computer-readable storage medium may be packaged together with the processor of the panoramic image generation device, or it may be packaged separately from the processor of the panoramic image generation device; this invention does not impose any limitations on this.
[0023] The descriptions of the second, third, fourth, and fifth aspects of this invention can be referred to the detailed description of the first aspect; and the beneficial effects described in the second, third, fourth, and fifth aspects can be referred to the analysis of the beneficial effects of the first aspect, which will not be repeated here.
[0024] In this invention, the name of the panoramic image generation device does not limit the device or functional module itself. In actual implementation, these devices or functional modules may appear under other names. As long as the function of each device or functional module is similar to that of this invention, it falls within the scope of the claims of this invention and its equivalents.
[0025] These or other aspects of the invention will become more apparent from the following description. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1a A schematic flowchart illustrating a method for generating panoramic images according to an embodiment of the present invention;
[0028] Figure 1b A top view schematic diagram of the layout of a ship environmental perception sensor provided in an embodiment of the present invention;
[0029] Figure 2 A schematic flowchart illustrating another method for generating panoramic images provided in an embodiment of the present invention;
[0030] Figure 3 A schematic diagram of the structure of a panoramic image generation device provided in an embodiment of the present invention;
[0031] Figure 4 This is a structural schematic diagram of a ship provided as an embodiment of the present invention. Detailed Implementation
[0032] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.
[0033] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0034] The terms "first" and "second," etc., used in the specification and drawings of this invention are used to distinguish different objects or to distinguish different treatments of the same object, rather than to describe a specific order of objects.
[0035] Furthermore, the terms "comprising" and "having," and any variations thereof, used in the description of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.
[0036] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but may also have additional steps not included in the figures. The process can correspond to a method, function, procedure, subroutine, subroutine, etc. Moreover, without conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0037] It should be noted that in the embodiments of the present invention, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of the present invention should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0038] In the description of this invention, unless otherwise stated, "a plurality of" means two or more.
[0039] Figure 1a This is a flowchart illustrating a panoramic image generation method according to an embodiment of the present invention. This embodiment is applicable to situations requiring the generation of panoramic images of a ship's operating area (i.e., its surrounding environment). The method can be applied to ships equipped with a lidar and dual-light pod on their top deck, and millimeter-wave radars on their bow and stern. The method can be executed by a panoramic image generation device, which can be implemented in software and / or hardware. For example, the device can be integrated into the ship's electronic equipment. (Reference) Figure 1a The panoramic image generation method in this embodiment specifically includes the following steps:
[0040] Step 110: Acquire point cloud data of the ship's operating area using lidar, acquire dual-light image data of the ship's operating area from multiple perspectives using a dual-light pod, and acquire obstacle detection data of the ship's operating area using millimeter-wave radar.
[0041] Specifically, a ship is a vehicle capable of navigating on water for operations or transportation. The ship's operational area is the spatial range centered on the ship when it performs its operational tasks (such as navigation). A lidar is a radar device that acquires three-dimensional spatial information of a target area by emitting laser light and receiving echo signals. A dual-light pod is an integrated camera pod device combining a visible light imaging module and an infrared thermal imaging module, capable of simultaneously acquiring visible light and infrared images of the target area. In this embodiment, it is used to acquire dual-light image data from multiple viewpoints of the ship, achieving multispectral and multi-view imaging of the operational area. The bow is the front part of the ship, i.e., the end of the hull facing the direction of travel when the ship is sailing. The stern is the rear part of the ship, opposite the bow, located at the stern in the direction of travel. A millimeter-wave radar is a radar device operating in the millimeter-wave frequency band, used to detect, locate, and identify obstacles within the target area to obtain obstacle detection data. Point cloud data is a set of discrete points acquired by the lidar, used to characterize the three-dimensional spatial coordinates of the target area (i.e., the ship's operational area), reflecting the three-dimensional structure and contour information of the environment. Dual-light image data refers to the collective visible light and infrared image data acquired by a dual-light pod, providing multispectral visual information about the target area. Obstacle detection data refers to the detection results obtained by millimeter-wave radar, characterizing the position, distance, and orientation of obstacles within the operational area.
[0042] In practice, the entire ship's operating area can be scanned by a lidar installed on the top of the ship to obtain point cloud data of the operating area; at the same time, a dual-light pod installed on the top of the ship can be used to collect dual-light image data of the operating area under preset different partition views; and millimeter-wave radars installed at the bow and stern of the ship can be used to detect and obtain obstacle detection data within the operating area.
[0043] For example, Figure 1b This is a top view schematic diagram of the layout of a ship environmental perception sensor provided in an embodiment of the present invention, as shown below. Figure 1b As shown in the figure, the ship's hull shape is clearly marked with a black pentagonal outline as a layout reference, where 1 indicates the bow direction and 2 indicates the stern direction. The sensor layout is as follows: the black diamond symbol at the top represents the millimeter-wave radar installed at the bow, the black diamond symbol at the bottom represents the millimeter-wave radar installed at the stern, the black circular symbol in the middle represents the lidar installed at the top of the ship (such as the top of the central area of the hull), and the gray circular symbol in the middle represents the dual-light pod installed at the top of the ship. In this embodiment, a preferred arrangement is four sets of dual-light pods, which can be evenly arranged at 90-degree (°) azimuth intervals around the top of the central area of the hull, facing the bow, stern, port side, and starboard side respectively.
[0044] It should be noted that, to ensure strict temporal alignment of point cloud data, dual-light image data, and obstacle detection data, and to guarantee the accuracy and consistency of subsequent panoramic image generation, the LiDAR, dual-light pod, and millimeter-wave radar employ a unified hardware clock or synchronous trigger signal for collaborative data acquisition. Specifically, synchronous trigger commands are sent to each environmental perception sensor, enabling the three types of devices to initiate data acquisition within the same time window, and the acquired data is marked using a unified sampling frequency and timestamp rules.
[0045] In this embodiment, through the above steps, environmental information of the ship's operating area can be comprehensively obtained from multiple dimensions such as three-dimensional spatial structure, multispectral visual images, and obstacle location information. This enables multi-source, multi-modal, and all-round perception of the operating area, effectively making up for the shortcomings of a single sensor in imaging, ranging, or obstacle detection. It improves the integrity and reliability of environmental perception and provides a rich and accurate data foundation for the subsequent generation of high-precision and robust panoramic images.
[0046] Step 120: Perform depth extraction processing on the point cloud data to obtain pixel-level depth information, and perform coordinate transformation processing on the dual-light image data based on the pixel-level depth information to obtain target visual data.
[0047] Specifically, pixel-level depth information refers to depth value data that corresponds one-to-one with each pixel in an image. It is used to characterize the actual spatial distance from the target point corresponding to each pixel in the image to the acquisition device. In this embodiment, pixel-level depth information characterizes the actual spatial distance from the target point corresponding to each pixel in the dual-light image to the acquisition device. Target visual data refers to the data obtained after coordinate transformation processing, which may include pixel information of the dual-light image, the three-dimensional coordinates corresponding to the pixels, depth values, and other information.
[0048] In practice, the point cloud data can first be preprocessed using a random sampling consensus algorithm or a height threshold method to remove invalid point clouds (such as sea clutter, radar noise points, etc.). Next, for each point in the preprocessed point cloud data, its three-dimensional coordinates are converted to coordinates in the dual-light pod coordinate system using a pre-calibrated extrinsic matrix of the lidar and the dual-light pod. Then, using the intrinsic matrix of the dual-light pod, these coordinates are mapped to image pixel coordinates. Simultaneously, a corresponding point cloud depth value (i.e., the z-axis distance in the dual-light pod coordinate system) is assigned to each pixel coordinate, obtaining pixel-level depth information.
[0049] Next, distortion correction is performed on the dual-light image data to obtain target dual-light image data without geometric distortion. Based on the obtained pixel-level depth information, the intrinsic parameter matrix of the dual-light pod, and the pre-calibrated rotation and translation matrix from the camera coordinate system to the lidar coordinate system, each effective pixel in the target dual-light image data is back-projected to three-dimensional space to obtain target visual data in the lidar coordinate system (i.e., the ship's coordinate system).
[0050] In this embodiment, the above steps achieve precise spatial alignment between image data and point cloud data, providing a rigorous data foundation for the subsequent generation of high-precision panoramic images.
[0051] Step 130: Cluster the point cloud data to obtain obstacle point cloud clusters, and extract features from the target visual data to obtain visual target detection results.
[0052] Specifically, an obstacle point cloud cluster is a subset of point clouds corresponding to a single obstacle target, obtained by clustering point cloud data. Each point cloud cluster represents a physical obstacle (such as a fishing boat or a buoy), including information such as the target's 3D contour and centroid coordinates. The visual target detection result is the obstacle detection information obtained after feature extraction from the target's visual data, including visible light detection results (such as the category, detection box, and confidence level of the obstacle in the visible light image) and infrared detection results (such as the category, detection box, and confidence level of the obstacle in the infrared image).
[0053] In practice, Euclidean clustering can be performed on the point cloud data to obtain obstacle point cloud clusters. Specifically, a K-dimensional tree spatial index structure can be constructed based on the point cloud data first. Then, a neighborhood radius (e.g., 0.5 meters) is set as a clustering constraint. The K-dimensional tree is used to traverse each point in the point cloud data, searching for all neighboring points within the neighborhood radius and merging the point and these neighboring points into the same cluster. Next, the above neighborhood search and merging operation is recursively performed on each newly added point until the current cluster can no longer include new neighboring points, thus completing the generation of a cluster. Finally, based on the volume of each point cloud cluster, the number of points N contained, and the filtering threshold, noisy clusters with too few points (e.g., N≤50) or too small a volume (e.g., less than 0.1 cubic meters) are removed to obtain obstacle point cloud clusters that conform to the characteristics of obstacles.
[0054] Simultaneously, the dual-light image from the target visual data is input into a pre-trained target detection model to obtain the visual target detection result. The pre-trained target detection model is a model obtained by supervised training of a deep learning model (such as YOLOv8, Faster R-CNN, etc.) based on historical dual-light image samples and corresponding obstacle annotation information in the ship operation scenario.
[0055] In this embodiment, the above steps improve the accuracy and effectiveness of obstacle localization, clarify the specific shape and visual attributes of visual targets, achieve accurate classification and image-level localization of obstacles, make up for the limitations of single-modal visual detection, and provide reliable data support for the subsequent generation of panoramic images.
[0056] Step 140: Generate a current panoramic image of the ship's operating area based on obstacle detection data, obstacle point cloud clusters, and visual target detection results.
[0057] Specifically, the current panoramic image is a complete image covering the ship's operating area (such as 360° around the ship or a specific sector of interest) generated by fusing obstacle detection data, obstacle point cloud clusters, and visual target detection results.
[0058] In the specific implementation, firstly, panoramic stitching processing is performed on the dual-light image data from multiple viewing angles to generate a dual-light panoramic visual base map covering the entire ship operation area. Next, based on the water surface plane assumption or digital elevation model, the preprocessed obstacle detection data is mapped to the corresponding pixel positions on the base map, and radar target icons (such as red dots or rhombuses) are drawn at these locations, along with their relative speed, distance, and other motion information. Subsequently, based on the spatial region where the obstacle point cloud clusters are located, the corresponding dual-light pod viewpoint is selected, and the 3D bounding box of the obstacle point cloud clusters is projected onto the pixel plane of the dual-light panoramic visual base map using its intrinsic and extrinsic parameter matrices. A corresponding 2D projection is then drawn on the dual-light panoramic visual base map. Based on this, using the mapping relationship from dual-light image data to the dual-light panoramic visual base map, perspective transformation is performed on the 2D bounding box coordinates in the visual target detection results to ensure they accurately match the targets on the dual-light panoramic visual base map. Finally, above each target in the dual-light panoramic visual base map, the category label (such as "fishing vessel," "buoy," etc.) and confidence level from the visual detection results are labeled. After overlaying the above multi-source data, a panoramic image that fully presents the distribution and status of obstacles in the ship's operating area can be obtained.
[0059] It should be noted that before overlaying all data onto the dual-light panoramic visual base map, the three-dimensional coordinates under the lidar coordinate system (i.e., the ship's coordinate system) should be used as a unified benchmark to correlate and match the detection results of millimeter-wave radar targets, obstacle point cloud clusters, and visual targets. For detection results belonging to the same physical target, only one set of comprehensive annotations (such as fusing radar distance, motion information, and visual target detection results) should be retained to avoid duplicate annotations and information redundancy, thereby improving the simplicity and readability of the panoramic image.
[0060] In this embodiment, the above steps enable the precise overlay and fusion of obstacle detection data, obstacle point cloud clusters, and visual target detection results in the panoramic image, significantly improving the accuracy, reliability, and information richness of the generated panoramic image.
[0061] The panoramic image generation method provided in this invention first acquires point cloud data of the ship's operating area using LiDAR, then acquires dual-light image data of the ship's operating area from multiple perspectives using a dual-light pod, and acquires obstacle detection data of the ship's operating area using millimeter-wave radar. This method comprehensively acquires environmental information of the ship's operating area from multiple dimensions, including three-dimensional spatial structure, multispectral visual images, and obstacle location information, achieving multi-source, multi-modal, and omnidirectional perception of the operating area. It effectively compensates for the shortcomings of single sensors in imaging, ranging, or obstacle detection, improving the completeness and reliability of environmental perception and providing a rich and accurate data foundation for the subsequent generation of high-precision, robust panoramic images. Next, depth extraction processing is performed on the point cloud data to obtain pixel-level depth information. Based on this pixel-level depth information, coordinate transformation processing is performed on the dual-light image data to obtain target visual data. This achieves precise spatial alignment between the image data and the point cloud data, providing a rigorous registration data foundation for the subsequent generation of high-precision panoramic images. Subsequently, the point cloud data is clustered to obtain obstacle point cloud clusters, and feature extraction is performed on the target visual data to obtain visual target detection results. This improves the accuracy and effectiveness of obstacle localization, clarifies the specific shape and visual attributes of visual targets, and achieves precise obstacle classification and image-level localization. It overcomes the limitations of single-modal visual detection and provides reliable data support for the subsequent generation of panoramic images. Finally, based on the obstacle detection data, obstacle point cloud clusters, and visual target detection results, a current panoramic image of the ship's operating area is generated. This achieves precise overlay and fusion of obstacle detection data, obstacle point cloud clusters, and visual target detection results in the panoramic image, significantly improving the accuracy, reliability, and information richness of the generated panoramic image. Therefore, the technical solution of this invention solves the problem in the prior art where it is difficult to accurately determine the distance between the ship and obstacles and the three-dimensional shape of the obstacles, resulting in insufficient accuracy and reliability of the generated panoramic image.
[0062] Figure 2 This is a flowchart illustrating another method for generating panoramic images according to an embodiment of the present invention. This embodiment is a specific modification based on the above embodiments. In this embodiment, the method may further include:
[0063] Step 210: Obtain point cloud data of the ship's operating area using lidar, obtain dual-light image data of the ship's operating area from multiple perspectives using a dual-light pod, and obtain obstacle detection data of the ship's operating area using millimeter-wave radar.
[0064] Step 211: Perform depth extraction processing on the point cloud data to obtain pixel-level depth information, and perform coordinate transformation processing on the dual-light image data based on the pixel-level depth information to obtain target visual data.
[0065] Furthermore, the point cloud data is subjected to depth extraction processing to obtain pixel-level depth information, including: spatial projection processing of the point cloud data based on the intrinsic and extrinsic parameter matrices of the dual-light pod to obtain the pixel position information and vertical position information of the point cloud data in the dual-light pod coordinate system; and determining the pixel-level depth information of the pixel corresponding to the pixel position information based on the pixel position information and the vertical position information corresponding to the pixel position information.
[0066] Specifically, the intrinsic parameter matrix of the dual-light pod refers to the matrix used to describe the imaging parameters inside the dual-light pod, including parameters such as focal length, principal point coordinates, and pixel size. The extrinsic parameter matrix of the dual-light pod refers to the matrix used to describe the spatial position and attitude transformation relationship of the dual-light pod coordinate system relative to the lidar coordinate system, and is usually composed of a rotation matrix and a translation vector. The dual-light pod coordinate system refers to a right-handed coordinate system established with the optical center of the dual-light pod as the origin. For example, the Z-axis of the dual-light pod coordinate system points forward of the pod lens (i.e., the shooting direction), the X-axis points horizontally to the right, and the Y-axis points vertically downward. In this embodiment, the lidar coordinate system is a three-dimensional rectangular coordinate system established with the center of the ship as the origin, where the X-axis points horizontally to the right along the ship, the Y-axis points vertically upward, and the Z-axis points longitudinally forward along the ship. Pixel position information refers to the two-dimensional pixel coordinate information corresponding to the dual-light image obtained after spatial projection of point cloud data, including row and column positions. Vertical position information refers to the depth value of the projected point cloud in the dual-light pod coordinate system, that is, the vertical distance from the point to the imaging plane of the dual-light pod.
[0067] In practice, the point cloud data can first be transformed from the LiDAR coordinate system to the dual-light pod coordinate system using an extrinsic parameter matrix, obtaining the three-dimensional coordinates of the point cloud in the dual-light pod coordinate system, which includes vertical position information (i.e., depth value). Then, the three-dimensional points are projected onto the image plane (i.e., the plane where the dual-light image is located) using the intrinsic parameter matrix of the dual-light pod, obtaining the corresponding pixel position information. Finally, based on the pixel position and its corresponding vertical position information, the pixel-level depth information of each pixel is determined.
[0068] In this embodiment, the above steps achieve precise spatial registration and depth mapping between 3D point cloud data and 2D dual-light image data, ensuring strict spatial alignment between 3D and 2D data, thereby improving the accuracy and practicality of subsequent panoramic image generation.
[0069] Optionally, coordinate transformation processing can be performed on the dual-light image data based on pixel-level depth information to obtain the coordinates of the dual-light image data in the lidar coordinate system. This can be achieved through the following formula:
[0070] ;
[0071] in, The three-dimensional coordinates of the dual-light image data in the lidar coordinate system; These are the homogeneous pixel coordinates in the dual-light pod coordinate system. This refers to pixel position information. represents the pixel-level depth information for the corresponding pixel; K is the intrinsic parameter matrix of the dual-light pod; R is the rotation matrix; t is the translation vector.
[0072] Step 212: Cluster the point cloud data to obtain obstacle point cloud clusters, and extract features from the target visual data to obtain visual target detection results.
[0073] Optionally, visual target detection results include visible light detection results and infrared detection results.
[0074] Step 213: Perform distortion correction and stitching processing on the dual-light image data to obtain a dual-light panoramic map.
[0075] Specifically, visible light detection results refer to the detection results obtained after target detection based on visible light images, used to characterize the location, category, and confidence information of targets within the ship's operating area in the visible light image. Infrared detection results refer to the detection results obtained after target detection based on infrared images, used to characterize the location, category, and confidence information of targets within the ship's operating area in the infrared image. Dual-light panoramic maps refer to panoramic images that contain both visible light and infrared information, obtained after distortion correction and image stitching.
[0076] In practice, the intrinsic parameter matrix and distortion coefficients (such as radial distortion coefficient and tangential distortion coefficient) pre-calibrated by the dual-light pod are first used to correct the distortion of the dual-light image data to obtain the corrected dual-light image data; then, the corrected dual-light image data is processed by multi-view stitching to obtain a dual-light panoramic map.
[0077] In this embodiment, the geometric errors caused by lens imaging distortion can be eliminated through the above steps, and the multi-view images can be integrated into a complete and unified dual-light panoramic map, thereby improving the geometric accuracy and visual coherence of the subsequently generated panoramic images.
[0078] Optionally, the dual-light image data includes visible light image data and infrared image data, and the dual-light panoramic map includes a visible light panoramic map and an infrared panoramic map.
[0079] Further, step 213 may specifically include: performing distortion correction processing on the visible light image data from each viewpoint to obtain visible light image data of each target, and performing distortion correction processing on the infrared image data from each viewpoint to obtain infrared image data of each target; extracting the same-name feature points from the visible light image data of each target, and fusing and stitching the visible light image data of each target based on the same-name feature points to obtain a visible light panoramic map; and performing projection stitching processing on the infrared image data of each target to obtain an infrared panoramic map.
[0080] Specifically, visible light image data refers to image data obtained by the visible light imaging module (built-in visible light camera) in the dual-light pod, used to reflect the color, texture, and contour information of the ship's operating area. Infrared image data refers to image data obtained by the infrared imaging module (built-in infrared camera) in the dual-light pod, used to reflect the temperature distribution and target thermal radiation characteristics of the ship's operating area. Target visible light image data refers to visible light image data obtained by distortion correction processing. Target infrared image data refers to infrared image data obtained by distortion correction processing. Visible light panoramic map refers to a panoramic image obtained by distortion correction and fusion stitching of multi-view visible light images. Infrared panoramic map refers to a panoramic image obtained by distortion correction and projection stitching of multi-view infrared images. Corresponding feature points refer to feature points in images from different viewpoints that correspond to the same physical location in real space, used to achieve multi-view image registration and stitching.
[0081] In the specific implementation, using the pre-calibrated intrinsic parameter matrices and distortion coefficients of the visible light and infrared cameras, the visible light image data and infrared image data from each viewpoint are subjected to distortion correction processing through a pixel remapping algorithm to obtain the visible light image data and infrared image data of each target.
[0082] For the corrected visible light image data of each target, a robust feature extraction algorithm (such as scale-invariant feature transformation or fast robust feature extraction) is first used to extract key feature points and descriptors for each target visible light image. Then, a Euclidean distance matching strategy is used to filter out corresponding feature points between adjacent viewpoint images (i.e., feature points corresponding to the same physical location in space), and a random sampling consensus algorithm is used to remove mismatched point pairs, resulting in filtered corresponding feature point pairs. Next, the homography matrix of adjacent images is calculated based on the filtered corresponding feature points. Based on the homography matrix, a spatial transformation is performed on each target visible light image to obtain spatially aligned target visible light images. For the overlapping areas of the spatially aligned adjacent images, a weighted fade-in / fade-out fusion algorithm is used for fusion processing, and pixel completion and boundary optimization are performed on non-overlapping areas to obtain a visible light panoramic map.
[0083] For the corrected infrared image data of each target, using a preset panoramic projection plane as a reference, the pixels of the infrared images of each viewpoint are mapped to a unified lidar coordinate system through the extrinsic parameter matrix of the dual-light pod, and then projected onto the panoramic image coordinate system to obtain infrared images of each target from each viewpoint with unified coordinates. Based on the projected coordinate positions, the pixel information of the infrared images from each viewpoint is mapped to the panoramic canvas; overlapping areas are processed using a grayscale mean fusion strategy to remove invalid pixels generated by projection, and the stitching edges are smoothed to obtain an infrared panoramic map. The preset panoramic projection plane is a reference plane pre-set according to actual conditions or requirements for unified projection stitching of multi-view images. The panoramic image coordinate system is a two-dimensional rectangular coordinate system established with the upper left corner (or center) of the generated panoramic image as the origin and pixels as the unit, used to locate the pixel positions of each projected image.
[0084] In this embodiment, the accuracy of the subsequently obtained panoramic image is improved through the above steps.
[0085] Furthermore, before step 213, the method further includes: filtering the obstacle detection data based on a preset signal-to-noise ratio threshold and a preset reflection cross-sectional area threshold to obtain initial obstacle detection data; performing dynamic compensation processing on the initial obstacle detection data according to the ship's own speed and relative speed to obtain intermediate obstacle detection data; and performing coordinate transformation processing on the intermediate obstacle detection data to obtain updated obstacle detection data.
[0086] Specifically, the preset signal-to-noise ratio (SNR) threshold is a critical value for distinguishing between valid signals and noise, pre-set according to actual conditions or requirements. The preset reflection cross-sectional area (RCA) threshold is a critical value for determining whether a target is a valid obstacle, pre-set according to actual conditions or requirements. The initial obstacle detection data is obstacle detection data obtained after filtering. The ship's own speed is the ship's own sailing speed. The ship's relative speed is the speed of the obstacle relative to the ship. The intermediate obstacle detection data is obstacle detection data obtained after dynamic compensation processing. The updated obstacle detection data is accurate and consistent obstacle detection data obtained after filtering, dynamic compensation, and coordinate transformation.
[0087] In practice, obstacle detection data can first be parsed and standardized to obtain preprocessed obstacle detection data. Then, based on a preset signal-to-noise ratio (SNR) threshold... th ) and preset cross-sectional area threshold (RCS) th The preprocessed obstacle detection data is filtered to identify valid targets, resulting in initial obstacle detection data (Target). valid Its filtering logic is: Target valid=(SNR≥SNR th )∧(RCS≥RCS th Next, dynamic compensation processing is performed on the initial obstacle detection data: based on the ship's own speed (V). own ) and the relative velocity detected by radar (V) rel ), calculate the absolute velocity (v) of the target (i.e., the obstacle) relative to the ground. abs ), to obtain intermediate obstacle detection data, the calculation formula is: v abs =V rel +V own , where V own With v abs The same coordinate system and unit system are used. Finally, the intermediate obstacle detection data is transformed from the radar coordinate system to the lidar coordinate system to obtain the updated obstacle detection data. The transformation formula is: Pl = R1·Pr + t1, where Pl is the obstacle coordinate in the lidar coordinate system, R1 is the rotation matrix of the radar coordinate system relative to the lidar coordinate system, Pr is the obstacle coordinate in the radar coordinate system, and t1 is the translation vector of the radar coordinate system relative to the lidar coordinate system.
[0088] In this embodiment, the above steps can improve data reliability, eliminate the influence of the ship's motion, and achieve spatial alignment of multi-source detection data, providing a unified and accurate data foundation for the subsequent generation of panoramic images.
[0089] Step 214: Match the visible light detection results and the infrared detection results to obtain cross-modal homologous targets.
[0090] Specifically, cross-modal homologous targets are detection targets that correspond to the same object in real space in both visible light and infrared images.
[0091] In the specific implementation, the visible light detection results of the visible light image and the infrared detection results of the infrared image are obtained respectively, and the target bounding boxes corresponding to each detection result are extracted; the intersection-union ratio of the two target bounding boxes is calculated, and the visible light targets and infrared targets with an intersection-union ratio greater than or equal to a preset threshold are associated and matched to obtain cross-modal homogeneous targets.
[0092] In this embodiment, the above steps achieve precise alignment between visible light and infrared targets, thereby improving the completeness and reliability of target recognition and thus enhancing the accuracy of the subsequent panoramic image.
[0093] Step 215: Determine the obstacle detection data and / or obstacle point cloud clusters corresponding to the cross-modal homogeneous target based on the preset spatial threshold, and obtain the target feature set.
[0094] Specifically, the preset spatial threshold is a spatial distance critical value pre-set according to actual conditions or needs. The target feature set is a set of target features corresponding to cross-modal homologous targets, composed of obstacle detection data and / or obstacle point cloud clusters.
[0095] In the specific implementation, the spatial coordinate information of the cross-modal co-source target in the dual-light panoramic map (i.e., the three-dimensional spatial coordinates obtained by mapping the image pixel coordinates of the co-source target) is used. Then, the spatial location information contained in the obstacle detection data and obstacle point cloud clusters is extracted. Next, based on the aforementioned data, the spatial distance between the co-source target and each obstacle detection data and obstacle point cloud cluster is calculated. The calculated spatial distance is compared with a preset spatial threshold, and obstacle detection data and / or obstacle point cloud clusters with spatial distances less than or equal to the preset spatial threshold are selected, determining that they correspond to the current cross-modal co-source target. Finally, the selected corresponding data are integrated to obtain a set containing all relevant feature information (obstacle detection data and / or obstacle point cloud clusters) of the target, i.e., the target feature set.
[0096] In this embodiment, the above steps achieve accurate correlation and unified representation of the detection results of millimeter-wave radar, lidar and visual multimodal detection, providing standardized and high-quality data support for the subsequent fusion display, unified annotation and situational understanding of target information in panoramic images.
[0097] Furthermore, after step 215, the method further includes: determining whether there is any missing data; if there is missing data, performing a traversal and matching of missing data of different data types to obtain matching results; and determining target annotation information based on the matching results and the missing data.
[0098] Specifically, the missing data refers to obstacle detection data and / or obstacle point cloud clusters that were not included in the target feature set.
[0099] In the specific implementation, it is determined whether there is any missing data. If no missing data is found, step 216 is executed directly. If missing data is found, the missing data of different data types is traversed and matched: the spatial distance between the obstacle detection data and the obstacle point cloud clusters in the missing data is calculated one by one. If the distance is less than the matching threshold, the two are considered to be successfully matched and a matching result is formed; if the distance is greater than or equal to the temporary matching threshold, the matching is considered to have failed. Then, for the data pairs of successfully matched obstacle detection data and obstacle point cloud clusters, the features of the two are integrated to generate target annotation information; for the missing data that failed to match (isolated obstacle detection data or obstacle point cloud clusters), target annotation information is generated separately based on its own features.
[0100] In this embodiment, the above steps can prevent the omission of unrelated obstacle data and ensure the integrity of target annotation information.
[0101] Step 216: Concatenate the cross-modal homologous targets with their corresponding target feature sets to obtain target annotation information.
[0102] Specifically, the target annotation information is the annotation data obtained by feature concatenation, which contains multimodal information of the target.
[0103] In practice, modal visual features (such as visible / infrared detection category, spatial location, and confidence level) of cross-modal co-source targets are extracted, and obstacle physical features (such as velocity, distance, reflective cross-section, and point cloud contour) are extracted from the target feature set corresponding to the cross-modal co-source target. Then, the two types of features are concatenated according to a preset structured format to obtain a complete feature set containing "modal visual features + obstacle physical features," i.e., target annotation information. Furthermore, if no corresponding target feature set is found for a cross-modal co-source target, target annotation information is directly generated based on the modal visual features of that cross-modal co-source target, ensuring that all cross-modal co-source targets have corresponding target annotation information.
[0104] In this embodiment, the above steps enable the target annotation information to simultaneously possess multimodal visual and physical attribute dimensions, thereby enhancing the richness and practicality of the annotation information and providing comprehensive target data support for subsequent panoramic image overlay.
[0105] Step 217: Overlay the target annotation information onto the corresponding pixel position of the dual-light panoramic map to obtain the current panoramic image of the ship operation area.
[0106] In the specific implementation, firstly, the three-dimensional spatial coordinates of the target contained in the target annotation information are extracted and projected onto the two-dimensional pixel coordinate systems corresponding to the visible light panoramic map and the infrared panoramic map, respectively, to calculate the pixel coordinates of the target on the two types of panoramic maps. Then, based on the pixel coordinates, a target anchor frame of uniform specification is drawn at the corresponding position on the visible light panoramic map and the infrared panoramic map to ensure that the anchor frame position of the same target in the cross-modal map is accurately aligned. Finally, the key attribute data of the target (such as distance value, velocity vector arrow, target category label, etc.) are visualized and annotated next to the anchor frame, and the annotated visible light panoramic map and the infrared panoramic map are superimposed to form the current panoramic image of the ship operation area.
[0107] In this embodiment, through the above steps, the generated panoramic image can retain the original visual features of visible light and infrared light, and can intuitively display the core physical properties of the target, thereby improving the reliability, accuracy, information richness and practicality of the panoramic image, and making it easier for staff to quickly identify the target status in the ship's operating area.
[0108] Furthermore, before step 217, the method further includes: determining the current predicted position information of each historical obstacle based on the position information and motion vector information of each historical obstacle in the previous panoramic image; matching the target position information in the target annotation information with the predicted position information of each historical obstacle to obtain a position matching result; and updating the target identification information in the target annotation information where the position matching result is a successful match to the identification information of the corresponding historical obstacle in the position matching result.
[0109] Specifically, historical obstacles are labeled and identifiable obstacle targets from the previous panoramic image, including their position, motion state, and unique identifier. Motion vector information consists of vector data (such as velocity value, motion direction angle, acceleration, etc.) describing the obstacle's displacement direction, velocity magnitude, and trajectory trend within a unit of time. Current predicted position information is the estimated spatial position of the obstacle at the current moment, calculated using a kinematic model (such as a uniform velocity / uniform acceleration model) based on the historical obstacle's previous position and motion vector. Position matching result is obtained by comparing the current target's labeled position with the predicted positions of historical obstacles. Target identification information is coded information used to uniquely identify the current target and is the core identifier distinguishing different targets.
[0110] In practice, based on a preset kinematic prediction model (such as a uniform linear motion model) and the time interval between the previous frame and the current frame, the current predicted position of each historical obstacle is calculated. Then, according to the target position in the current target annotation information, the spatial Euclidean distance between this position and the current predicted position of each historical obstacle is calculated one by one, and the calculated distance is compared with a preset position matching threshold: if the distance is less than or equal to the threshold, the target is determined to have successfully matched the corresponding historical obstacle; otherwise, the match is determined to have failed. This yields the position matching result (including successfully matched target-historical obstacle pairs and newly matched targets that failed to match).
[0111] Next, for target annotation information with a successful location matching result, the unique identifier information of the corresponding historical obstacle is extracted; and the original target identifier information in the current target annotation information is updated to the identifier information of the historical obstacle; for target annotation information with a failed match, the original newly generated identifier information is retained as the unique identifier of the newly added obstacle.
[0112] In this embodiment, the above steps ensure that the identification information of the same obstacle remains consistent in different frames of panoramic images; avoid the same obstacle being repeatedly marked as a new target due to changes in position between frames, facilitate staff to trace the movement trajectory of obstacles, and further enhance the practical value of panoramic images.
[0113] The panoramic image generation method provided in this invention first acquires point cloud data of the ship's operating area using LiDAR, then acquires dual-light image data of the ship's operating area from multiple perspectives using a dual-light pod, and acquires obstacle detection data of the ship's operating area using millimeter-wave radar. This method comprehensively acquires environmental information of the ship's operating area from multiple dimensions, including three-dimensional spatial structure, multispectral visual images, and obstacle location information, achieving multi-source, multi-modal, and omnidirectional perception of the operating area. It effectively compensates for the shortcomings of single sensors in imaging, ranging, or obstacle detection, improving the completeness and reliability of environmental perception and providing a rich and accurate data foundation for the subsequent generation of high-precision, robust panoramic images. Next, depth extraction processing is performed on the point cloud data to obtain pixel-level depth information. Based on this pixel-level depth information, coordinate transformation processing is performed on the dual-light image data to obtain target visual data. This achieves precise spatial alignment between the image data and the point cloud data, providing a rigorous registration data foundation for the subsequent generation of high-precision panoramic images. Subsequently, the point cloud data was clustered to obtain obstacle point cloud clusters, and feature extraction was performed on the target visual data to obtain visual target detection results. This improved the accuracy and effectiveness of obstacle localization, clarified the specific shape and visual attributes of visual targets, and achieved precise obstacle classification and image-level localization, overcoming the limitations of single-modal visual detection and providing reliable data support for the subsequent generation of panoramic images. Then, distortion correction and stitching processing were performed on the dual-light image data to obtain a dual-light panoramic map. This eliminated geometric errors caused by lens imaging distortion and integrated multi-view images into a complete and unified dual-light panoramic map, improving the geometric accuracy and visual coherence of the subsequently generated panoramic images. Matching processing was performed on the visible light and infrared detection results to obtain cross-modal homogeneous targets, achieving precise correlation and alignment between visible light and infrared targets, thereby improving the completeness and reliability of target recognition and ultimately enhancing the accuracy of the subsequently obtained panoramic images. Based on a preset spatial threshold, obstacle detection data and / or obstacle point cloud clusters corresponding to cross-modal co-origin targets are determined, resulting in a target feature set. This achieves accurate correlation and unified representation of millimeter-wave radar, lidar, and visual multimodal detection results, providing standardized and high-quality data support for the fusion display, unified annotation, and situational understanding of target information in subsequent panoramic images. By concatenating the cross-modal co-origin targets with their corresponding target feature sets, target annotation information is obtained. This imbues the target annotation information with information from both multimodal visual and physical attribute dimensions, enhancing the richness and practicality of the annotation information and providing comprehensive target data support for subsequent panoramic image overlay.Finally, the target annotation information is superimposed onto the corresponding pixel positions of the dual-light panoramic map to obtain the current panoramic image of the ship's operating area. This ensures that the generated panoramic image retains the original visual features of visible light and infrared light while intuitively displaying the core physical properties of the target, thus improving the reliability, accuracy, information richness, and practicality of the panoramic image. This facilitates rapid identification of target status within the ship's operating area by personnel. Therefore, the technical solution of this invention solves the problem in existing technologies where it is difficult to accurately determine the distance between the ship and obstacles, as well as the three-dimensional shape of the obstacles, leading to insufficient accuracy and reliability in the generated panoramic images.
[0114] Figure 3 This is a schematic diagram of a panoramic image generation device provided in an embodiment of the present invention. The device is applied to a ship, which is equipped with a lidar and a dual-light pod on its top, and millimeter-wave radars on its bow and stern. This device and the panoramic image generation methods in the above embodiments belong to the same inventive concept. For details not described in detail in the embodiments of the panoramic image generation device, please refer to the embodiments of the panoramic image generation methods described above.
[0115] like Figure 3 As shown, the device includes:
[0116] The acquisition module 310 is used to acquire point cloud data of the ship's operating area through the lidar, acquire dual-light image data of the ship's operating area from multiple perspectives through the dual-light pod, and acquire obstacle detection data of the ship's operating area through the millimeter-wave radar.
[0117] The conversion module 320 is used to perform depth extraction processing on the point cloud data to obtain pixel-level depth information, and to perform coordinate transformation processing on the dual-light image data based on the pixel-level depth information to obtain target visual data.
[0118] The extraction module 330 is used to perform clustering processing on the point cloud data to obtain obstacle point cloud clusters, and to extract features from the target visual data to obtain visual target detection results;
[0119] The generation module 340 is used to generate a current panoramic image of the ship operation area based on the obstacle detection data, the obstacle point cloud clusters, and the visual target detection results.
[0120] Based on the above embodiments, the visual target detection result includes visible light detection result and infrared detection result. The generation module 340 is specifically used for:
[0121] The dual-light image data is subjected to distortion correction and stitching processing to obtain a dual-light panoramic map; the visible light detection results and the infrared detection results are matched to obtain cross-modal homogeneous targets; obstacle detection data and / or obstacle point cloud clusters corresponding to the cross-modal homogeneous targets are determined based on a preset spatial threshold to obtain a target feature set; the cross-modal homogeneous targets and the corresponding target feature set are stitched together to obtain target annotation information; the target annotation information is superimposed on the corresponding pixel positions of the dual-light panoramic map to obtain the current panoramic image of the ship operation area.
[0122] Based on the above embodiments, the device further includes:
[0123] The judgment module is used to determine whether there is any missing data after obtaining the target feature set by determining the obstacle detection data and / or obstacle point cloud clusters corresponding to the cross-modal homogeneous target based on a preset spatial threshold. The missing data refers to obstacle detection data and / or obstacle point cloud clusters that are not included in the target feature set. If there is missing data, the missing data of different data types are traversed and matched to obtain the matching result. The target annotation information is determined based on the matching result and the missing data.
[0124] Based on the above embodiments, the device further includes:
[0125] The matching module is used to determine the current predicted position information of each historical obstacle based on the position information and motion vector information of each historical obstacle in the previous panoramic image before overlaying the target annotation information onto the corresponding pixel position of the dual-light panoramic map to obtain the current panoramic image of the ship operation area; match the target position information in the target annotation information with the predicted position information of each historical obstacle to obtain a position matching result; and update the target identification information in the target annotation information that has a successful position matching result to the identification information of the corresponding historical obstacle in the position matching result.
[0126] Based on the above embodiments, the dual-light image data includes visible light image data and infrared image data, and the dual-light panoramic map includes a visible light panoramic map and an infrared panoramic map. The generation module 340 performs distortion correction and stitching processing on the dual-light image data to obtain a dual-light panoramic map, including:
[0127] Distortion correction processing is performed on visible light image data from each viewpoint to obtain visible light image data of each target, and distortion correction processing is performed on infrared image data from each viewpoint to obtain infrared image data of each target; corresponding feature points are extracted from the visible light image data of each target, and the visible light image data of each target are fused and stitched based on the corresponding feature points to obtain the visible light panoramic map; projection stitching processing is performed on the infrared image data of each target to obtain the infrared panoramic map.
[0128] Based on the above embodiments, the conversion module 320 is specifically used for:
[0129] Based on the intrinsic and extrinsic parameter matrices of the dual-light pod, the point cloud data is spatially projected to obtain the pixel position information and vertical position information of the point cloud data in the dual-light pod coordinate system; based on the pixel position information and the corresponding vertical position information, the pixel-level depth information of the pixel corresponding to the pixel position information is determined.
[0130] Based on the above embodiments, the device further includes:
[0131] The processing module is configured to, before generating a current panoramic image of the ship's operating area based on the obstacle detection data, the obstacle point cloud clusters, and the visual target detection results, filter the obstacle detection data based on a preset signal-to-noise ratio threshold and a preset reflectance cross-sectional area threshold to obtain initial obstacle detection data; perform dynamic compensation processing on the initial obstacle detection data based on the ship's own speed and relative speed to obtain intermediate obstacle detection data; and perform coordinate transformation processing on the intermediate obstacle detection data to obtain updated obstacle detection data.
[0132] The panoramic image generation apparatus provided in this embodiment of the invention can execute the panoramic image generation method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.
[0133] It is worth noting that in the embodiments of the above-mentioned panoramic image generation device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.
[0134] Figure 4 This is a structural schematic diagram of a ship provided as an embodiment of the present invention. Figure 4 A block diagram of an exemplary vessel 4 suitable for implementing embodiments of the present invention is shown. Figure 4 The ship 4 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0135] like Figure 4 As shown, the vessel 4 is represented in the form of a general-purpose computing electronic device. The components of the vessel 4 may include, but are not limited to: one or more processors or processing units 16, system memory 28, and a bus 18 connecting different system components (including system memory 28 and processing unit 16).
[0136] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0137] Ship 4 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by Ship 4, including volatile and non-volatile media, movable and non-movable media.
[0138] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Ship 4 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (… Figure 4 Not shown; usually referred to as a "hard drive"). Although Figure 4 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. System memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0139] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of the present invention.
[0140] Ship 4 can also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), and with one or more devices that enable a user to interact with Ship 4, and / or with any device that enables Ship 4 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed through input / output (I / O) interface 22. Furthermore, Ship 4 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 20. Figure 4 As shown, network adapter 20 communicates with other modules of vessel 4 via bus 18. It should be understood that, although... Figure 4 As not shown in the diagram, other hardware and / or software modules may be used in conjunction with Ship 4, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0141] The processing unit 16 executes various functional applications and page displays by running programs stored in the system memory 28, such as implementing the panoramic image generation method provided in the embodiments of the present invention.
[0142] Of course, those skilled in the art will understand that the processor can also implement the technical solution of the panoramic image generation method provided in any embodiment of the present invention.
[0143] This invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements, for example, the panoramic image generation method provided in this invention.
[0144] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0145] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0146] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0147] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0148] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computing device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0149] Furthermore, the acquisition, storage, use, and processing of data in the technical solution of this invention all comply with relevant laws and regulations.
[0150] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A method for generating panoramic images, characterized in that, Applied to ships, the ship is equipped with a lidar and a dual-light pod on its top, and millimeter-wave radars are installed on its bow and stern, respectively. The method includes: The laser radar acquires point cloud data of the ship's operating area, the dual-light pod acquires dual-light image data of the ship's operating area from multiple perspectives, and the millimeter-wave radar acquires obstacle detection data of the ship's operating area. The point cloud data is subjected to depth extraction processing to obtain pixel-level depth information, and the dual-light image data is subjected to coordinate transformation processing based on the pixel-level depth information to obtain target visual data. Clustering is performed on the point cloud data to obtain obstacle point cloud clusters, and feature extraction is performed on the target visual data to obtain visual target detection results; Based on the obstacle detection data, the obstacle point cloud clusters, and the visual target detection results, a current panoramic image of the ship's operating area is generated.
2. The method according to claim 1, characterized in that, The visual target detection results include visible light detection results and infrared detection results; Based on the obstacle detection data, the obstacle point cloud clusters, and the visual target detection results, a current panoramic image of the ship's operating area is generated, including: The dual-light image data is subjected to distortion correction and stitching processing to obtain a dual-light panoramic map; The visible light detection results and the infrared detection results are matched to obtain cross-modal homogeneous targets; Based on a preset spatial threshold, the obstacle detection data and / or obstacle point cloud clusters corresponding to the cross-modal homogeneous target are determined to obtain the target feature set; The cross-modal homogeneous targets are concatenated with their corresponding target feature sets to obtain target annotation information; The target annotation information is superimposed onto the corresponding pixel position of the dual-light panoramic map to obtain the current panoramic image of the ship operation area.
3. The method according to claim 2, characterized in that, After determining the obstacle detection data and / or obstacle point cloud clusters corresponding to the cross-modal homogeneous target based on a preset spatial threshold to obtain the target feature set, the method further includes: Determine whether there is any missing data, wherein the missing data is obstacle detection data and / or obstacle point cloud clusters that are not included in the target feature set; If there is missing data, the missing data of different data types are traversed and matched to obtain the matching results; the target annotation information is determined based on the matching results and the missing data.
4. The method according to claim 2, characterized in that, Before overlaying the target annotation information onto the corresponding pixel position of the dual-light panoramic map to obtain the current panoramic image of the ship operation area, the method further includes: Based on the position information and motion vector information of each historical obstacle in the previous panoramic image, the current predicted position information of each historical obstacle is determined. The target location information in the target annotation information is matched with the predicted location information of each historical obstacle to obtain the location matching result; Update the target identifier information in the target annotation information of the location matching result that is a successful match to the identifier information of the corresponding historical obstacle in the location matching result.
5. The method according to claim 2, characterized in that, The dual-light image data includes visible light image data and infrared image data, and the dual-light panoramic map includes a visible light panoramic map and an infrared panoramic map. Distortion correction and stitching processing are performed on the dual-light image data to obtain the dual-light panoramic map, including: Distortion correction processing is performed on the visible light image data from each viewpoint to obtain the visible light image data of each target, and distortion correction processing is performed on the infrared image data from each viewpoint to obtain the infrared image data of each target. Extract the corresponding feature points from the visible light image data of each target, and fuse and stitch the visible light image data of each target based on the corresponding feature points to obtain the visible light panoramic map; The infrared image data of each target are projected and stitched together to obtain the infrared panoramic map.
6. The method according to claim 1, characterized in that, The point cloud data is subjected to depth extraction processing to obtain pixel-level depth information, including: Based on the intrinsic and extrinsic parameter matrices of the dual-light pod, the point cloud data is subjected to spatial projection processing to obtain the pixel position information and vertical position information of the point cloud data in the dual-light pod coordinate system; Based on the pixel position information and the vertical position information corresponding to the pixel position information, the pixel-level depth information of the pixel corresponding to the pixel position information is determined.
7. The method according to claim 1, characterized in that, Before generating the current panoramic image of the ship's operating area based on the obstacle detection data, the obstacle point cloud clusters, and the visual target detection results, the method further includes: The obstacle detection data is filtered based on a preset signal-to-noise ratio threshold and a preset reflection cross-sectional area threshold to obtain initial obstacle detection data; The initial obstacle detection data is dynamically compensated based on the ship's own speed and relative speed to obtain intermediate obstacle detection data. The intermediate obstacle detection data is subjected to coordinate transformation to obtain updated obstacle detection data.
8. A panoramic image generation apparatus, characterized in that, Applied to ships, the ship is equipped with a lidar and a dual-light pod mounted on its top, and millimeter-wave radars are mounted on its bow and stern, respectively. The device includes: The acquisition module is used to acquire point cloud data of the ship's operating area through the lidar, acquire dual-light image data of the ship's operating area from multiple perspectives through the dual-light pod, and acquire obstacle detection data of the ship's operating area through the millimeter-wave radar. The conversion module is used to perform depth extraction processing on the point cloud data to obtain pixel-level depth information, and to perform coordinate transformation processing on the dual-light image data based on the pixel-level depth information to obtain target visual data. The extraction module is used to perform clustering processing on the point cloud data to obtain obstacle point cloud clusters, and to extract features from the target visual data to obtain visual target detection results; The generation module is used to generate a current panoramic image of the ship's operating area based on the obstacle detection data, the obstacle point cloud clusters, and the visual target detection results.
9. A ship, characterized in that, The vessels include: At least one processor; and a memory communicatively connected to said at least one processor; The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the panoramic image generation method according to any one of claims 1-7.
10. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the panoramic image generation method according to any one of claims 1-7.
Citation Information
Patent Citations
Multi-view road intelligent identification method based on dual-light fusion
CN113643345A
Ship berthing collision early warning method and system and storage medium
CN116110255A
Target tracking method based on visible light, infrared and laser radar data fusion
CN116258744A
Airport target detection method and system based on millimeter wave radar, laser radar and high-definition array camera
CN116977806A
Environment sensing method, early warning method, equipment, system and medium
CN117636276A