Target collaborative sensing method and system based on intelligent fleet navigation
Through the fusion of AIS and lidar and real-time processing of video data, combined with the screening of self-attention mechanism model and feature extraction of PCA analysis method, the problem of insufficient perception ability of intelligent ships in complex water environments is solved, efficient target recognition and collaborative perception are achieved, and the comprehensive environmental perception ability of multiple intelligent ships is improved.
Patent Information
- Application Number
- CN202510513242.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-23
AI Technical Summary
In complex water environments, the perception ability of a single intelligent ship is limited, making it difficult to obtain the surrounding target information comprehensively and accurately, resulting in a small perception range, insufficient information fusion, low target recognition accuracy, low comprehensive recognition ability, easy to form data islands, and difficult to improve the comprehensive operation efficiency of multiple intelligent ships.
Through the integration of AIS information and lidar detection data, the target space information is positioned and perceived; the video data of multiple intelligent ships is obtained in real time, keyframes are extracted, and the object feature set is constructed using Sobel operator and PCA analysis method; the data transmission volume is evaluated in combination with network status, and the self-attention mechanism model is used to filter high-scoring features, target detection and recognition are carried out to achieve coordinated target perception.
It significantly improves the target recognition accuracy and system response speed, rationally utilizes multi-intelligent ship data for target collaborative perception, effectively improving the comprehensive environmental perception ability of multiple ships.
Smart Images

Figure CN120047877A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent fleets, and more specifically, to a method and system for collaborative target perception based on the navigation of intelligent fleets. Background Art
[0002] With the continuous development of fisheries and shipping, intelligent fleets are playing an increasingly important role in fields such as marine environmental monitoring, resource exploration, search and rescue missions, and water area fishing. However, in a complex water area environment, the perception ability of a single ship is limited, making it difficult to comprehensively and accurately obtain surrounding target information and achieve efficient collaborative operations. In traditional intelligent ship navigation and task execution, there is a lack of comprehensive target recognition methods for multi-source data, multi-sensor data, and multi-video source data. Traditional methods are often limited to object recognition of a single ship, suffering from problems such as a small perception range, insufficient information fusion, and low target recognition accuracy. Their comprehensive recognition ability is low, prone to forming data islands, which is not conducive to collaborative environment perception analysis of multiple intelligent ships and difficult to improve the comprehensive operation efficiency of multiple intelligent ships.
[0003] Therefore, there is an urgent need for a method for collaborative target perception based on the navigation of intelligent fleets to solve the above problems. Summary of the Invention
[0004] The present invention overcomes the defects of the prior art and provides a method and system for collaborative target perception based on the navigation of intelligent fleets.
[0005] The first aspect of the present invention provides a method for collaborative target perception based on the navigation of intelligent fleets, including: S102: In multiple intelligent ships, locate and perceive target spatial information through AIS information and lidar detection data; S104: Based on multiple intelligent ships, obtain multiple video data based on target perception in real time; S106: Extract key frames from multiple video data, segment the object regions of the key frames based on the sobel operator, and extract the main components of object features through PCA analysis to form multiple first feature sets; S108: Evaluate the real-time shared data volume through the intelligent ship network status; S110: Train a self-attention mechanism model through the feature map data of the recognized objects, determine the query matrix, key matrix, and value matrix, combine every two first feature sets, calculate the attention scores of the combined feature sets based on the self-attention mechanism model, screen the features with high scores to obtain the features that meet the real-time shared data volume, and mark them as the second feature sets; S112: Based on the second feature set, perform target detection and recognition through an image recognition model to generate a recognition result. Based on the recognition result and the target space information, perform target collaborative perception, and send the perception result to multiple intelligent ships.
[0006] In this solution, the S102 is specifically as follows: The AIS information includes the azimuth, heading, and speed information of each intelligent ship. Perform spatial analysis on the lidar detection data to generate point cloud data, perform filtering processing on the point cloud data, and segment different target objects to obtain the target object bounding box based on the point cloud information. Perform spatial deviation averaging of the target object based on the target object bounding boxes analyzed by multiple intelligent ships to obtain the initial spatial information based on the point cloud information. Combined with the AIS information, import the position information of each intelligent ship into the initial spatial information to generate the target space information.
[0007] In this solution, the S104 is specifically as follows: Through the video monitoring device, among multiple intelligent ships, perceive the surrounding target objects and perform real-time video shooting on the target range to obtain the corresponding video data. Store the video data in the intelligent ship database and share the data through the Internet of Things.
[0008] In this solution, the S106 is specifically as follows: Extract key frames from the video data to obtain key image frames. Perform preprocessing on the key image frames, including image enhancement, image noise reduction, and normalization. Based on the sobel operator, perform gradient calculation based on the gray value for each pixel point in the key image frame, divide the edges through the gradient value, and perform object area segmentation through the edge division to obtain the object area. Based on the object area, extract color and edge features in the key image frame. During the process of dividing the edges through the gradient value for the edge features, perform edge feature extraction, and obtain the color features through color histogram statistics. Based on the PCA analysis method, use the color and edge features as inputs and generate two input feature vectors. Matrixize the input feature vectors to form a feature matrix, and calculate the covariance matrix C of the feature matrix. Perform eigenvalue decomposition on the covariance matrix C to obtain the decomposed eigenvalues and decomposed eigenvectors. Sort the decomposed eigenvalues by size and extract the top 70% of the decomposed eigenvectors as the principal components. Principal component extraction is performed based on two input feature vectors to obtain the first feature set.
[0009] In this solution, S108 is specifically: Within a real-time period, the packet loss rate, network latency, and network jitter of each intelligent ship are statistically calculated. Based on the packet loss rate, network latency, and network jitter, mean calculation is performed to obtain the homogenized network parameters. Through the homogenized network parameters, the maximum data transmission volume within a real-time period is evaluated to obtain the real-time shared data volume.
[0010] In this solution, S110 includes: Obtain the feature map data and feature recognition rate of the existing recognized objects from the system database. Set the attention score based on the feature recognition rate. Associate the feature map data with the attention score. Construct a self-attention mechanism model, and through a linear transformation method, perform mapping transformation on the feature map data to generate an initial query matrix, an initial key matrix, and an initial value matrix. Import the feature map data into the self-attention mechanism model and calculate the attention score corresponding to the feature map. Use the absolute difference between the attention score and the attention score as the loss function to train the self-attention mechanism model, and repeatedly adjust the query matrix, the initial key matrix, and the initial value matrix until the loss function converges to a preset difference. After training is completed, determine the query matrix, the key matrix, and the value matrix.
[0011] In this solution, S110 further includes: Combine every two first feature sets to obtain a combined feature set. Import the combined feature set into the self-attention mechanism model, calculate the attention score for the feature vectors therein, and screen out the feature sets with high scores, which are marked as high-score feature data. Perform high-score feature data analysis and screening on all combined feature sets, and based on the real-time shared data volume, select the feature data that meets the data volume from the high-score feature data, which is marked as the second feature set.
[0012] In this solution, S112 is specifically: Import the second feature set into the image recognition model for object classification and object image positioning, and generate a recognition result. Based on the recognition result and the target space information, perform target collaborative perception, and send the perception result to multiple intelligent ships through the central platform.
[0013] In a second aspect of the present invention, there is also provided a target collaborative perception system based on intelligent fleet navigation, the system comprising: a memory, a processor, wherein the memory includes a target collaborative perception program based on intelligent fleet navigation, and when the target collaborative perception program based on intelligent fleet navigation is executed by the processor, the following steps are implemented: S102: Among multiple intelligent ships, locate and perceive target space information through AIS information and lidar detection data; S104: Based on multiple intelligent ships, obtain multiple video data based on target perception in real time; S106: Extract key frames from multiple video data, segment the object regions of the key frames based on the sobel operator, and extract the main components of the object features through PCA analysis to form multiple first feature sets; S108: Evaluate the real-time shared data volume through the intelligent ship network state; S110: Train a self-attention mechanism model through the feature map data of the recognized objects, determine the query matrix, key matrix, and value matrix, combine every two first feature sets, calculate the attention scores of the combined feature sets based on the self-attention mechanism model, screen the features with high scores, obtain the features that meet the real-time shared data volume, and mark them as the second feature sets; S112: Through an image recognition model, perform target detection and recognition based on the second feature sets, generate recognition results, perform target collaborative perception based on the recognition results and target space information, and send the perception results to multiple intelligent ships.
[0014] In a third aspect of the present invention, there is also provided a computer-readable storage medium, wherein the computer-readable storage medium includes a target collaborative perception program based on intelligent fleet navigation, and when the target collaborative perception program based on intelligent fleet navigation is executed by a processor, the steps of the target collaborative perception method based on intelligent fleet navigation as described in any one of the above are implemented.
[0015] The present invention discloses a target collaborative perception method and system based on intelligent fleet navigation. The target space information is obtained through the fusion positioning of AIS and lidar, and multi-view video streams are collected in real time based on multiple intelligent ships. After the system extracts the video key frames, Sobel edge detection and PCA analysis techniques are used to construct an object feature set, and the data transmission volume is determined by combining network state dynamic evaluation. The pre-trained self-attention model is used to perform correlation analysis on the feature sets, and the features with high attention scores are screened out to achieve efficient target recognition and shared perception. The present invention can significantly improve the target recognition accuracy and system response speed, rationally utilize the data of multiple intelligent ships for target collaborative perception, and effectively improve the comprehensive environment perception ability of multiple ships. Description of the Drawings
[0016] Figure 1 The flowchart of a method for collaborative target perception based on intelligent fleet navigation according to the present invention is shown; Figure 2 The block diagram of a system for collaborative target perception based on intelligent fleet navigation according to the present invention is shown. Specific embodiments
[0017] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.
[0018] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0019] Figure 1 The flowchart of a method for collaborative target perception based on intelligent fleet navigation according to the present invention is shown.
[0020] As Figure 1 shown, the first aspect of the present invention provides a method for collaborative target perception based on intelligent fleet navigation, including: S102: In multiple intelligent ships, locate and perceive the target space information through AIS information and lidar detection data; S104: Based on multiple intelligent ships, obtain multiple video data based on target perception in real time; S106: Extract key frames from multiple video data, segment the object regions of the key frames based on the sobel operator, and extract the main components of the object features through PCA analysis to form multiple first feature sets; S108: Evaluate the real-time shared data volume through the intelligent ship network status; S110: Train a self-attention mechanism model through the feature map data of the recognized objects, determine the query matrix, key matrix and value matrix, combine every two first feature sets, calculate the attention scores of the combined feature sets based on the self-attention mechanism model, and screen the features with high scores to obtain the features that meet the real-time shared data volume, which are marked as the second feature sets; S112: Through an image recognition model, perform target detection and recognition based on the second feature sets, generate recognition results, perform collaborative target perception based on the recognition results and the target space information, and send the perception results to multiple intelligent ships.
[0021] According to an embodiment of the present invention, the S102 is specifically: The AIS information includes the azimuth, heading, and speed information of each intelligent ship; Perform spatial analysis on the lidar detection data to generate point cloud data, filter the point cloud data, and segment different target objects to obtain the target object bounding box based on the point cloud information; Perform spatial deviation averaging of the target objects based on the target object bounding boxes analyzed from multiple intelligent ships to obtain the initial spatial information based on the point cloud information; Combine the AIS information, import the position information of each intelligent ship into the initial spatial information, and generate the target spatial information.
[0022] It should be noted that in the target perception, generally, intelligent recognition and perception are performed on objects such as other ships, water area shore buildings, buoys, docks, etc., so that the intelligent ship can better complete tasks such as water area monitoring, surveillance, search and rescue, and transportation.
[0023] According to an embodiment of the present invention, the S104 is specifically: Through the video monitoring device, among multiple intelligent ships, perceive the surrounding target objects and perform real-time video shooting on the target range to obtain the corresponding video data; Store the video data in the intelligent ship database and share the data through the Internet of Things.
[0024] It should be noted that for the data sharing through the Internet of Things, multiple intelligent ships and the central platform can share data through the Internet of Things to achieve collaborative perception of target objects.
[0025] According to an embodiment of the present invention, the S106 is specifically: Extract key frames from the video data to obtain key image frames; Perform image enhancement, image noise reduction, and normalization preprocessing on the key image frames; Based on the sobel operator, perform gradient calculation based on the gray value for each pixel point in the key image frame, divide the edges through the gradient value, and segment the object area through the edge division to obtain the object area; Based on the object area, extract color and edge features in the key image frame; During the process of dividing the edges through the gradient value for the edge features, extract the edge features, and obtain the color features through color histogram statistics; Based on the PCA analysis method, use the color and edge features as inputs and generate two input feature vectors; Matrixize the input feature vectors to form a feature matrix, and calculate the covariance matrix C of the feature matrix; Perform eigenvalue decomposition on the covariance matrix C to obtain the decomposed eigenvalues and decomposed eigenvectors; Sort the decomposed eigenvalues by magnitude and extract the top 70% of the decomposed eigenvectors as the principal components; Extract the principal components based on the two input eigenvectors to obtain the first feature set.
[0026] It should be noted that the Sobel operator filters the edges by calculating the gray-scale gradient of each pixel point to adapt to the effective separation of object images in the water environment. Additionally, based on the analysis requirements, other edge detection algorithms can be used for object region segmentation. Before calculating the gradient of the gray-scale value, image graying is required. The two input eigenvectors respectively represent the corresponding eigenvectors of color features and edge features, and both vectors include multiple ones. The top 70% refers to the top 70% of the highest-score data. Multiple video data generate corresponding multiple first feature sets. The first feature set includes the principal component features of both color and edge data.
[0027] According to an embodiment of the present invention, the S108 is specifically: Within a real-time period, count the packet loss rate, network latency, and network jitter of each intelligent ship; Calculate the average value based on the packet loss rate, network latency, and network jitter to obtain the homogenized network parameters; Evaluate the maximum data transmission volume within a real-time period through the homogenized network parameters to obtain the real-time shared data volume.
[0028] It should be noted that the homogenized network parameters include the average packet loss rate, average network latency, and average network jitter. Homogenization means averaging the network parameters of multiple intelligent ships.
[0029] According to an embodiment of the present invention, the S110 includes: Obtain the feature map data and feature recognition rate of the existing recognized objects from the system database; Set the attention score based on the feature recognition rate; Associate the feature map data with the attention score; Construct a self-attention mechanism model, and through a linear transformation method, perform a mapping transformation on the feature map data to generate an initial query matrix, an initial key matrix, and an initial value matrix; Import the feature map data into the self-attention mechanism model and calculate the attention score corresponding to the feature map. Use the absolute difference between the attention score and the attention score as the loss function to train the self-attention mechanism model, and cyclically adjust the query matrix, the initial key matrix, and the initial value matrix until the loss function converges to a preset difference; After training is completed, determine the query matrix, key matrix, and value matrix.
[0030] It should be noted that the feature map data is the feature map obtained during the recognition process in the historical recognized object image database. Generally speaking, when performing image recognition based on a CNN network or other artificial intelligence recognition models for image analysis, relevant feature maps usually need to be extracted during the intermediate recognition process, and the feature maps can be collected and stored based on the platform database. The feature recognition rate is the successful recognition rate after object recognition using the relevant feature maps, which can reflect to a certain extent the importance of the feature in the actual recognition process. In the feature map data, each feature map includes an attention score. The attention score is directly proportional to the feature recognition rate. For example, if the feature recognition rate is 50%, the attention score can be set to the same value, 50 points.
[0031] The loss function can be set as f = |attention score - attention score value|. By minimizing the difference between the two, the self-attention model is trained to obtain the learning parameters (query matrix, key matrix, and value matrix). The preset difference can be set to 10, that is, the score difference is within 10.
[0032] When importing the feature map data into the self-attention mechanism model and calculating the attention score corresponding to the feature map, specifically through the calculation of the score matrix Z, the corresponding score is obtained, and the calculation is as follows: , where Z is the finally calculated score matrix, and the attention score is obtained through this matrix. Q, K, V, are the dimensions of the query matrix, key matrix, value matrix, and vectors in the feature map respectively.
[0033] According to the embodiment of the present invention, the S110 further includes: Combining every two first feature sets to obtain a combined feature set; Importing the combined feature set into the self-attention mechanism model, calculating the attention score for the feature vectors therein, and screening out the feature sets with high scores, which are marked as high-score feature data; Performing high-score feature data analysis and screening on all combined feature sets, and based on the real-time shared data volume, selecting the feature data that meets the data volume from the high-score feature data, which is marked as the second feature set.
[0034] It should be noted that if the high-score feature data volume is greater than the real-time shared data volume, random screening can be performed to reduce a certain amount of feature data to meet the data volume. Additionally, it is also possible to sort and screen the higher-score features based on the score. The role of combination is to be able to fuse the principal component feature data obtained by different intelligent ships and perform fusion screening on the feature data to obtain high-score features. If there are three first feature sets, the number of pairwise combinations is three, and three combined feature sets can be obtained through three combinations.
[0035] According to an embodiment of the present invention, the S112 is specifically as follows: Import the second feature set into the image recognition model for object classification and object image positioning, and generate a recognition result; Based on the recognition result and the target space information, perform target collaborative perception, and send the perception result to multiple intelligent ships through the central platform.
[0036] It should be noted that the image recognition model can be a model that utilizes relevant image feature recognition, such as yolov5, an image recognition model based on CNN, a support vector machine, a decision tree, etc., which are models for image recognition and classification.
[0037] In the present invention, the entire operation system of intelligent ships includes multiple intelligent ships and a central platform. The central platform can send the perception analysis result to multiple intelligent ships. Additionally, based on the requirements set by the terminal, a certain intelligent ship can be used as the central platform for target collaborative perception. The first feature set and the second feature set both include two-dimensional features of color and edge.
[0038] Here, it is worth mentioning that in the traditional navigation and task execution of intelligent ships, there is a lack of a comprehensive target recognition method for multi-source data, multi-sensor data, and multi-video source data. The traditional method is often limited to the object recognition of a single ship, with low comprehensive recognition ability, and it is unable to effectively combine the multi-source data of other nearby intelligent ships for shared analysis or collaborative perception, easily forming data islands, which is not conducive to the collaborative environment perception analysis of multiple intelligent ships and difficult to improve the comprehensive operation efficiency of multiple intelligent ships. Based on this, the present invention obtains the key frames of the video through PCA analysis for the video acquisition of intelligent ships, and performs a first feature extraction on the key frames of the video. This feature extraction is used to obtain the principal component features for target recognition. In the second step, a self-attention mechanism model is trained to calculate the attention scores for different feature data using the self-attention mechanism model, and calculate the scores in the principal component features (i.e., the first feature set). Through the score situation, the second feature set with high scores is selected. Based on the shared network among intelligent ships, the fusion analysis of the source video data is realized. After obtaining the second feature set, the target can be collaboratively perceived and efficiently recognized, and the recognition result is combined with the spatial information of the target perception, effectively realizing precise target perception and spatial confirmation, and effectively improving the comprehensive environment perception ability of multiple ships.
[0039] Based on the shared network, multiple intelligent ships can share the corresponding target perception results.
[0040] According to an embodiment of the present invention, it further includes: Obtain the attention scores of the feature data based on the second feature set; Generate multiple second feature maps according to the feature vectors in the second feature set, and associate the second feature maps with the attention scores; According to the DCA feature fusion algorithm, multiply multiple second feature maps with the first weight respectively to obtain multiple class activation maps; Perform spatially weighted fusion on the multiple class activation maps. During the weighted fusion process, each class activation map corresponds to a scoring weight value, and the scoring weight value is equal to the corresponding attention score, and finally obtain the fused feature data; Perform object detection and recognition on the fused feature data.
[0041] It should be noted that during the image recognition process, since the second feature set is data for obtaining image features from multiple intelligent ships in multi-angle video data, the data has multi-dimensionality and multi-source characteristics. When a general image recognition model performs feature analysis, there may be a situation of low feature processing efficiency for this feature set. Therefore, in the present invention, by improving the DCA fusion method, the feature data (i.e., feature vectors) in the second feature set are feature-fused, and the attention score is used as the weight for weighted fusion during the fusion process to improve the recognition efficiency and feature processing efficiency of the fused features in the subsequent process, greatly improving the adaptability to various image recognition models, and further improving the practicality and efficiency of the target perception system.
[0042] The first weight is a user preset value. One second feature map corresponds to one class activation map and corresponds to one attention score (also used as the scoring weight value).
[0043] According to an embodiment of the present invention, forming multiple first feature sets further includes: Judge whether the quantity of the first feature sets is greater than a preset multiple of the real-time shared data quantity. If so, construct an autoencoder; Based on the object region, extract color and edge features in the key image frame, and import the color and edge features as the original input features into the autoencoder. Map the input features to low-dimensional representation data through the encoder; Based on the decoder, reconstruct the low-dimensional representation data and generate reconstructed feature data; Based on the mean square error as the loss function, evaluate the difference between the reconstructed feature data and the original input features, and cyclically optimize the autoencoder until the loss function converges to the expected value, and vectorize the reconstructed feature data at this time to obtain the first feature set.
[0044] It should be noted that the preset multiple can be set to 2 or more, which is used to determine whether the data volume of the first feature set is much larger than the real-time shared data volume. If so, it means that the feature data volume extracted from the current multiple video data is large, the dimension is high, and the redundancy is high. It is difficult to reduce the data volume through the PCA extraction method, and it is also necessary to ensure the retention of the main feature data. Therefore, when the data volume is much larger than the real-time shared data volume, the present invention performs corresponding dimensionality reduction operations on the feature data through an autoencoder to further reduce data redundancy, extract the feature data of the main dimension, and form the first feature set. In the determination of the first feature set, that is, based on the feature set obtained in advance by the PCA method, if the data volume of the first feature set is much larger than the real-time shared data volume, then re-analysis is performed based on the color and edge features to generate a new first feature set. The expected value is a smaller result value of the loss function, which can be set by the user.
[0045] Figure 2 Fig. shows a block diagram of an object collaborative perception system based on intelligent fleet navigation according to the present invention.
[0046] In a second aspect of the present invention, there is also provided an object collaborative perception system 2 based on intelligent fleet navigation. The system includes: a memory 21 and a processor 22. The memory 21 includes an object collaborative perception program based on intelligent fleet navigation. When the object collaborative perception program based on intelligent fleet navigation is executed by the processor 22, the following steps are implemented: S102: In multiple intelligent ships, locate and perceive the target space information through AIS information and lidar detection data; S104: Based on multiple intelligent ships, obtain multiple video data based on target perception in real time; S106: Extract key frames from multiple video data, segment the object areas of the key frames based on the sobel operator, and extract the main components of the object features through PCA analysis to form multiple first feature sets; S108: Evaluate the real-time shared data volume through the intelligent ship network status; S110: Train a self-attention mechanism model through the feature map data of the recognized objects, determine the query matrix, key matrix, and value matrix, combine every two first feature sets, calculate the attention score of the combined feature set based on the self-attention mechanism model, and screen the features with high scores to obtain the features that meet the real-time shared data volume, which are marked as the second feature set; S112: Through the image recognition model, perform target detection and recognition based on the second feature set, generate a recognition result, perform object collaborative perception based on the recognition result and the target space information, and send the perception result to multiple intelligent ships.
[0047] According to an embodiment of the present invention, the S102 is specifically: The AIS information includes the azimuth, course, and speed information of each intelligent ship; Perform spatial analysis on the lidar detection data to generate point cloud data, filter the point cloud data, and segment different target objects to obtain the target object bounding box based on the point cloud information; Perform spatial deviation averaging of the target object based on the target object bounding boxes analyzed from multiple intelligent ships to obtain the initial spatial information based on the point cloud information; Combine the AIS information, import the position information of each intelligent ship into the initial spatial information, and generate the target spatial information.
[0048] It should be noted that in the target perception, generally, intelligent recognition and perception of other ships, water area shore buildings, buoys, docks and other objects are performed to enable the intelligent ship to better complete tasks such as water area monitoring, surveillance, search and rescue, and transportation.
[0049] According to an embodiment of the present invention, the S104 is specifically: Through the video monitoring device, among multiple intelligent ships, perceive the surrounding target objects and perform real-time video shooting on the target range to obtain the corresponding video data; Store the video data in the intelligent ship database and share the data through the Internet of Things.
[0050] It should be noted that for the data sharing through the Internet of Things, multiple intelligent ships and the central platform can share data through the Internet of Things to achieve collaborative perception of target objects.
[0051] According to an embodiment of the present invention, the S106 is specifically: Extract key frames from the video data to obtain key image frames; Perform image enhancement, image noise reduction, and normalization preprocessing on the key image frames; Based on the sobel operator, perform gradient calculation based on the gray value for each pixel point in the key image frame, divide the edges through the gradient value, and segment the object area through the edge division to obtain the object area; Based on the object area, extract color and edge features in the key image frame; During the process of dividing the edges through the gradient value for the edge features, extract the edge features, and obtain the color features through color histogram statistics; Based on the PCA analysis method, use the color and edge features as inputs and generate two input feature vectors; Matrixize the input feature vectors to form a feature matrix, and calculate the covariance matrix C of the feature matrix; Perform eigenvalue decomposition on the covariance matrix C to obtain the decomposed eigenvalues and eigenvectors. Sort the decomposed eigenvalues by magnitude and extract the top 70% of the decomposed eigenvectors as the principal components. Extract principal components based on two input eigenvectors to obtain the first feature set.
[0052] It should be noted that the Sobel operator screens edges by calculating the gray gradient of each pixel point to adapt to the effective separation of object images in the water environment. Additionally, based on the analysis requirements, other edge detection algorithms can be used for object area segmentation. Before calculating the gradient of the gray value, image grayscaling is required. The two input eigenvectors respectively represent the corresponding eigenvectors of color features and edge features, and both vectors include multiple ones. The top 70% refers to the top 70% of the highest-scoring data. Multiple video data generate corresponding multiple first feature sets. The first feature set includes the principal component features of both color and edge data.
[0053] According to an embodiment of the present invention, the S108 is specifically as follows: Within a real-time period, count the packet loss rate, network latency, and network jitter of each intelligent ship. Calculate the average value based on the packet loss rate, network latency, and network jitter to obtain the homogenized network parameters. Evaluate the maximum data transmission volume within a real-time period through the homogenized network parameters to obtain the real-time shared data volume.
[0054] It should be noted that the homogenized network parameters include the average packet loss rate, average network latency, and average network jitter. Homogenization means averaging the network parameters of multiple intelligent ships.
[0055] According to an embodiment of the present invention, the S110 includes: Obtain the feature map data and feature recognition rate of the recognized objects from the system database. Set the attention score based on the feature recognition rate. Associate the feature map data with the attention score. Construct a self-attention mechanism model, and through a linear transformation method, perform mapping transformation on the feature map data to generate an initial query matrix, an initial key matrix, and an initial value matrix. Import the feature map data into the self-attention mechanism model and calculate the attention score corresponding to the feature map. Use the absolute difference between the attention score and the attention score as the loss function to train the self-attention mechanism model, and repeatedly adjust the query matrix, the initial key matrix, and the initial value matrix until the loss function converges to a preset difference. After training is completed, determine the query matrix, key matrix, and value matrix.
[0056] It should be noted that the feature map data is the feature map obtained during the recognition process in the historical recognized object image database. Generally speaking, when performing image recognition based on a CNN network or other artificial intelligence recognition models for image analysis, relevant feature maps usually need to be extracted during the intermediate recognition process, and the feature maps can be collected and stored based on the platform database. The feature recognition rate is the successful recognition rate after object recognition using the relevant feature maps, and it can reflect to a certain extent the importance of the feature in the actual recognition process. In the feature map data, each feature map includes an attention score. The attention score is directly proportional to the feature recognition rate. For example, if the feature recognition rate is 50%, the attention score can be set to the same value, 50 points.
[0057] The loss function can be set as f = |attention score - attention degree score|. By minimizing the difference between the two, the self-attention model is trained to obtain the learning parameters (query matrix, key matrix, and value matrix). The preset difference can be set to 10, that is, the score difference is within 10.
[0058] When importing the feature map data into the self-attention mechanism model and calculating the attention score corresponding to the feature map, specifically, through the calculation of the score matrix Z, the corresponding score value is obtained, and the calculation is as follows: , where Z is the finally calculated score matrix, and the attention score value is obtained through this matrix. Q, K, V, are the query matrix, key matrix, value matrix, and the dimensionality of the vectors in the feature map, respectively.
[0059] According to the embodiment of the present invention, the S110 further includes: Combining every two first feature sets to obtain a combined feature set; Importing the combined feature set into the self-attention mechanism model, calculating the attention scores for the feature vectors therein, and screening out the feature sets with high scores, which are marked as high-score feature data; Performing high-score feature data analysis and screening on all combined feature sets, and based on the real-time shared data volume, selecting the feature data that meets the data volume from the high-score feature data, which is marked as the second feature set.
[0060] It should be noted that if the high-score feature data volume is greater than the real-time shared data volume, random screening can be performed to reduce a certain amount of feature data to meet the data volume. Additionally, higher-score features can also be screened based on the score ranking. The role of combination is to be able to fuse the principal component feature data obtained by different intelligent ships and perform fusion screening on the feature data to obtain high-score features. If there are three first feature sets, the number of pairwise combinations is three, and three combined feature sets can be obtained through three combinations.
[0061] According to an embodiment of the present invention, the S112 is specifically as follows: Import the second feature set into the image recognition model for object classification and object image positioning, and generate a recognition result; Based on the recognition result and the target space information, perform target collaborative perception, and send the perception result to multiple intelligent ships through the central platform.
[0062] It should be noted that the image recognition model can be a model that utilizes relevant image feature recognition, such as yolov5, an image recognition model based on CNN, a support vector machine, a decision tree, and other models for image recognition and classification.
[0063] In the present invention, the entire operation system of the intelligent ship includes multiple intelligent ships and a central platform. The central platform can send the perception analysis result to multiple intelligent ships. Additionally, based on the requirements set by the terminal, a certain intelligent ship can be used as the central platform for target collaborative perception. The first feature set and the second feature set both include two-dimensional features of color and edge.
[0064] It is worth mentioning here that in the traditional navigation and task execution of intelligent ships, there is a lack of a comprehensive target recognition method for multi-source data, multi-sensor data, and multi-video source data. The traditional method is often limited to the object recognition of a single ship, with low comprehensive recognition ability, and it is unable to effectively combine the multi-source data of other nearby intelligent ships for shared analysis or collaborative perception, easily forming data islands, which is not conducive to the collaborative environment perception analysis of multiple intelligent ships and difficult to improve the comprehensive operation efficiency of multiple intelligent ships. Based on this, the present invention obtains the video of the intelligent ship, and based on the PCA analysis method, performs a primary feature extraction on the key frames of the video. This feature extraction is used to obtain the principal component features for target recognition. In the second step, a self-attention mechanism model is trained to calculate the attention scores for different feature data and calculate the scores in the principal component features (i.e., the first feature set). Through the score situation, the second feature set with high scores is selected. Based on the shared network among intelligent ships, the fusion analysis of the source video data is realized. After obtaining the second feature set, the target can be collaboratively perceived and efficiently recognized, and the recognition result is combined with the spatial information of the target perception to effectively achieve precise target perception and space confirmation, effectively improving the comprehensive environment perception ability of multiple ships.
[0065] Based on the shared network, multiple intelligent ships can share the corresponding target perception results.
[0066] The third aspect of the present invention further provides a computer-readable storage medium, which includes a target collaborative perception program based on intelligent fleet navigation. When the target collaborative perception program based on intelligent fleet navigation is executed by a processor, the steps of the target collaborative perception method based on intelligent fleet navigation as described in any one of the above are implemented.
[0067] The present invention discloses a target collaborative perception method and system based on intelligent fleet navigation. The target spatial information is obtained by fusing the positioning of AIS and lidar, and multi-view video streams are collected in real time based on multiple intelligent ships. After the system extracts the key frames of the video, the Sobel edge detection and PCA analysis techniques are used to construct an object feature set, and the data transmission volume is determined by combining the dynamic evaluation of the network state. The pre-trained self-attention model is used to perform correlation analysis on the feature set, and the features with high attention scores are screened out to achieve efficient target recognition and shared perception. The present invention can significantly improve the target recognition accuracy and the system response speed, reasonably utilize the data of multiple intelligent ships for target collaborative perception, and effectively improve the comprehensive environment perception ability of multiple ships.
[0068] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed may be through some interfaces, and the indirect coupling or communication connection of the devices or units may be electrical, mechanical, or other forms.
[0069] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they may be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0070] In addition, each functional unit in the embodiments of the present invention can be all integrated in one processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit; the above integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0071] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments; and the foregoing storage medium includes: various media that can store program codes, such as removable storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0072] Alternatively, if the above integrated units are implemented in the form of software function modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. And the foregoing storage medium includes: various media that can store program codes, such as removable storage devices, ROM, RAM, magnetic disks, or optical discs.
[0073] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention.
Claims
1. A target collaborative perception method based on intelligent fleet navigation, characterized in that: include: S102: Among multiple intelligent ships, positioning and sensing target space information through AIS information and laser radar detection data; S104: Based on multiple intelligent ships, multiple video data based on target perception are acquired in real time; S106: extracting key frames from the multiple video data, segmenting the key frames into object regions based on the Sobel operator, and extracting principal components of object features through a PCA analysis method to form multiple first feature sets; S108: Real-time shared data volume through intelligent ship network status assessment; S110: training a self-attention mechanism model through feature graph data of existing recognized objects, and determining a query matrix, a key matrix, and a value matrix, combining every two first feature sets, and calculating the attention score of the combined feature set based on the self-attention mechanism model, screening with high-scoring features, obtaining features that meet the real-time shared data volume, and marking them as the second feature set; S112: Target detection and recognition are performed based on the second feature set through the image recognition model to generate a recognition result, target collaborative perception is performed based on the recognition result and target spatial information, and the perception result is sent to multiple smart ships.
2. The target collaborative perception method based on intelligent fleet navigation according to claim 1 is characterized in that: The S102 is specifically: The AIS information includes the position, heading and speed information of each smart ship; Perform spatial analysis through LiDAR detection data to generate point cloud data, filter the point cloud data, segment different target objects, and obtain the target object bounding box based on point cloud information; The spatial deviation of the target object is averaged according to the target object bounding box obtained by analyzing multiple intelligent ships to obtain the initial spatial information based on the point cloud information; Combined with AIS information, the location information of each smart ship is imported into the initial spatial information to generate target spatial information.
3. The target collaborative perception method based on intelligent fleet navigation according to claim 1 is characterized in that: The S104 is specifically: Through the video monitoring device, multiple intelligent ships can sense surrounding target objects and take real-time video of the target range to obtain corresponding video data; The video data is stored in the smart ship database and shared through the Internet of Things.
4. The target collaborative perception method based on intelligent fleet navigation according to claim 1 is characterized in that: The S106 is specifically: Extract key frames from video data to obtain key image frames; Perform image enhancement, image noise reduction, and standardization preprocessing on key image frames; Based on the Sobel operator, the gradient of each pixel in the key image frame is calculated based on the gray value, and the edge is divided by the gradient value. The object area is segmented by edge division to obtain the object area; Based on the object area, color and edge feature extraction is performed in the key image frame; Edge features are extracted during the process of dividing edges by gradient values, and color features are obtained through color histogram statistics; Based on the PCA analysis method, color and edge features are used as input and two input feature vectors are generated; Matrix the input feature vector to form a feature matrix, and calculate the covariance matrix C of the feature matrix; Perform eigenvalue decomposition on the covariance matrix C to obtain decomposition eigenvalues and decomposition eigenvectors; Sort the decomposed eigenvalues by size and extract the first 70% of the decomposed eigenvectors as the principal components; Principal components are extracted based on the two input feature vectors to obtain a first feature set.
5. The target collaborative perception method based on intelligent fleet navigation according to claim 1 is characterized in that: The S108 is specifically: In a real-time cycle, the packet loss rate, network delay and network jitter of each smart ship are counted; Based on the packet loss rate, network delay and network jitter, the average value is calculated to obtain the averaged network parameters; The maximum data transmission volume in a real-time cycle is evaluated by averaging network parameters to obtain the real-time shared data volume.
6. The target collaborative perception method based on intelligent fleet navigation according to claim 1 is characterized in that: The S110 includes: Obtain feature map data and feature recognition rate of identified objects from the system database; Setting attention scores based on feature recognition rates; Associating feature graph data with attention scores; Construct a self-attention mechanism model, map and transform the feature map data through linear transformation, and generate the initial query matrix, initial key matrix and initial value matrix; Import the feature map data into the self-attention mechanism model and calculate the attention score corresponding to the feature map. Use the absolute difference between the attention score and the attention score as the loss function to train the self-attention mechanism model, and cyclically adjust the query matrix, initial key matrix, and initial value matrix until the loss function converges to the preset difference. After training is completed, the query matrix, key matrix and value matrix are determined.
7. The target collaborative perception method based on intelligent fleet navigation according to claim 6 is characterized in that: The S110 further includes: Combining every two first feature sets to obtain a combined feature set; Import the combined feature set into the self-attention mechanism model, calculate the attention score of the feature vectors therein, and filter out the feature sets with high scores and mark them as high-score feature data; The high-scoring feature data of all combined feature sets are analyzed and screened, and based on the real-time shared data volume, feature data that meets the data volume is selected from the high-scoring feature data and marked as obtaining the second feature set.
8. The target collaborative perception method based on intelligent fleet navigation according to claim 7 is characterized in that: The S112 is specifically: Importing the second feature set into the image recognition model to perform object classification and object image positioning, and generate a recognition result; Based on the recognition results and target space information, target collaborative perception is performed, and the perception results are sent to multiple smart ships through the central platform.
9. A target collaborative perception system based on intelligent fleet navigation, characterized in that: The system includes: a memory and a processor. The memory includes a target collaborative perception program based on intelligent fleet navigation. When the target collaborative perception program based on intelligent fleet navigation is executed by the processor, the following steps are implemented: S102: Among multiple intelligent ships, positioning and sensing target space information through AIS information and laser radar detection data; S104: Based on multiple intelligent ships, multiple video data based on target perception are acquired in real time; S106: extracting key frames from the multiple video data, segmenting the key frames into object regions based on the Sobel operator, and extracting principal components of object features through a PCA analysis method to form multiple first feature sets; S108: Real-time shared data volume through intelligent ship network status assessment; S110: training a self-attention mechanism model through feature graph data of existing recognized objects, and determining a query matrix, a key matrix, and a value matrix, combining every two first feature sets, and calculating the attention score of the combined feature set based on the self-attention mechanism model, screening with high-scoring features, obtaining features that meet the real-time shared data volume, and marking them as the second feature set; S112: Target detection and recognition are performed based on the second feature set through the image recognition model to generate a recognition result, target collaborative perception is performed based on the recognition result and target spatial information, and the perception result is sent to multiple smart ships.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a target collaborative perception program based on intelligent fleet navigation. When the target collaborative perception program based on intelligent fleet navigation is executed by a processor, the steps of the target collaborative perception method based on intelligent fleet navigation as described in any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Multi-unmanned-aerial-vehicle and multi-unmanned-ship inspection control system based on reinforcement learning
CN113671994A
Unmanned intelligent multi-mode information fusion and target perception system and operation method
CN116310689A
Intelligent navigation safety management method and system based on multi-source data fusion
CN117975769A
Ship obstacle avoidance and automatic berthing method based on deep learning algorithm
CN119088034A
Intelligent sensing method for automatic driving of small ship
CN119693922A
Cited By
Intelligent ship communication network integrated monitoring method and device
CN120811928A
Channel buoy detection method based on fusion of multi-mode pulse neural network and visual Transform
CN121616952A