Positioning method, device and equipment based on feature screening in ship environment
By screening out features with high importance and stability in the ship environment and constructing a positioning model, the problem of low positioning accuracy in the ship environment is solved, and higher positioning accuracy and model adaptability are achieved.
Patent Information
- Application Number
- CN202510800607.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-10-03
AI Technical Summary
In a ship environment, traditional wireless positioning methods have low positioning accuracy due to factors such as electromagnetic reflection, refraction, and obstruction, and dynamic disturbance factors increase signal instability, making it difficult to ensure high-precision positioning.
By obtaining multiple sets of positioning-related data, a gradient decision tree model is constructed, the importance and stability of features are evaluated, the target feature set for positioning is screened out, a positioning model is constructed, and positioning is performed using the screened features.
It improves the positioning accuracy and the generalization ability of the model in different environments, reduces prediction fluctuations, and adapts to the resource-constrained edge computing needs in the ship environment.
Smart Images

Figure CN120751479A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of wireless positioning technology, and in particular to a positioning method, device and equipment based on feature screening in a ship environment. Background Art
[0002] With the continuous development of the Internet of Things and wireless communication technologies, indoor positioning technology is increasingly being used in industrial manufacturing, intelligent logistics, and ship management. In the unique environment of ships, indoor positioning can be used to achieve functions such as personnel positioning, safety management, cargo tracking, and intelligent inspections, with great practical value.
[0003] However, compared to conventional building environments, ship interiors exhibit significant structural differences and electromagnetic characteristics. First, the hull structure generally utilizes large areas of metal, which are highly electromagnetically reflective. Second, the enclosed interior space, densely packed cabins, and complex layout lead to strong reflections, refractions, and obstructions during wireless signal propagation, resulting in severe multipath effects and frequency drift. Furthermore, dynamic disturbances during navigation, such as deck vibrations and slight shifts in the relative position of walls, further exacerbate signal instability.
[0004] In the above complex context, traditional positioning methods based on received signal strength, delay or phase are difficult to guarantee high-precision positioning performance. Summary of the Invention
[0005] In view of this, the present application provides a positioning method, apparatus and device based on feature screening in a ship environment, for accurately performing positioning in a ship environment.
[0006] Specifically, this application is implemented through the following technical solutions:
[0007] A first aspect of the present application provides a method for feature-based screening in a ship environment, the method comprising:
[0008] Acquire multiple sets of positioning-related data; wherein each set of positioning-related data includes an IQ data sequence acquired by a receiving end in the process of receiving a constant tone extension (CTE) signal transmitted from a transmitter, and the location of the transmitter;
[0009] Using the IQ data sequence in each set of positioning-related data as an input feature and the position in the set of positioning-related data as a label, and using the multiple sets of positioning-related data to construct a first training sample set;
[0010] Training a gradient decision tree model using the first training sample set, and after the training is completed, determining an importance score for each feature in the input features based on the training results;
[0011] For each feature, the trained gradient decision tree model is used to evaluate the predictive ability of the feature under independent input conditions to obtain the overall stability score of the feature; wherein the overall stability score is used to indicate the stability of the feature in maintaining predictive accuracy under different positions;
[0012] Determining a comprehensive score for each feature according to the importance score and the stability score of each feature, and screening a target feature set for positioning from multiple features included in the input features based on the comprehensive score of each feature;
[0013] Based on the target feature set and the training sample set, construct a second training sample set for training a positioning model, and train the positioning model with the second training sample set;
[0014] After the data to be predicted is acquired, feature values corresponding to the target feature set are extracted from the data to be predicted, and the extracted feature values are input into the positioning model so that the positioning model predicts the position corresponding to the data to be predicted.
[0015] The second aspect of the present application provides a positioning device based on feature screening in a ship environment, the device comprising an acquisition module, a construction module, a calculation module, and a prediction module, wherein:
[0016] The acquisition module is configured to acquire multiple sets of positioning-related data; wherein each set of positioning-related data includes an IQ data sequence acquired by the receiving end in the process of receiving a constant tone extension (CTE) signal transmitted from a transmitter, and the location of the transmitter;
[0017] The construction module is configured to use the IQ data sequence in each set of positioning-related data as an input feature and the position in the set of positioning-related data as a label, and to construct a first training sample set using the multiple sets of positioning-related data;
[0018] The computing module is configured to train a gradient decision tree model using the first training sample set, and after the training is completed, determine an importance score for each feature in the input features based on the training results;
[0019] The calculation module is used to evaluate the predictive ability of each feature under independent input conditions using a trained gradient decision tree model to obtain an overall stability score for the feature; wherein the overall stability score is used to indicate the degree to which the feature maintains predictive accuracy at different positions;
[0020] The determination module is configured to determine a comprehensive score for each feature based on the importance score and stability score of each feature, and to select a target feature set for positioning from a plurality of features included in the input features based on the comprehensive score of each feature;
[0021] The construction module is configured to construct a second training sample set for training a positioning model based on the target feature set and the training sample set, and train the positioning model using the second training sample set;
[0022] The prediction module is used to extract feature values corresponding to the target feature set from the data to be predicted after obtaining the data to be predicted, and input the extracted feature values into the positioning model so that the positioning model predicts the position corresponding to the data to be predicted.
[0023] The third aspect of the present application provides a feature-based screening device in a ship environment, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps of any one of the methods provided in the first aspect of the present application are implemented.
[0024] The positioning method, device and equipment based on feature screening in the ship environment provided in this application eliminate irrelevant or unstable features by combining the feature importance score with the overall stability score, retain the key features with high discrimination for the positioning task, and effectively improve the accuracy of the predicted position. In addition, when determining the overall stability score of the feature, the consistency of the performance of the feature in different positions is comprehensively considered, which can improve the generalization ability of the model in different cabins and different environments and reduce prediction fluctuations; further, by reducing the input dimension through feature screening and reducing the training and reasoning burden of the model, the real-time performance of the system can be improved, and it can adapt to the resource-constrained edge computing deployment requirements in the ship environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 This is a flowchart of Example 1 of the positioning method based on feature screening in a ship environment provided by this application;
[0026] Figure 2 This is a structural diagram of a first embodiment of a positioning device based on feature screening in a ship environment provided by this application;
[0027] Figure 3 This is a hardware structure diagram of a positioning device based on feature screening in a ship environment where the positioning device based on feature screening in the ship environment of this application is located. DETAILED DESCRIPTION
[0028] Exemplary embodiments are described in detail herein, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numerals in different drawings represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with this application.
[0029] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a," "the," and "the" used in this application are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0030] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0031] Specific embodiments are given below to introduce the technical solutions of the present application in detail.
[0032] Figure 1 This is a flowchart of the first embodiment of the positioning method based on feature screening in a ship environment provided by this application. Figure 1 The method provided in this embodiment may include:
[0033] S101. Acquire multiple groups of positioning-related data; wherein each group of the positioning-related data includes an IQ data sequence acquired by a receiving end in the process of receiving a constant tone extension (CTE) signal transmitted from a transmitter, and the location of the transmitter.
[0034] Specifically, the transmitters emitting the constant tone extended CTE signal can be Bluetooth tags based on Bluetooth 5.1. These tags are installed within the ship's cabin and periodically transmit the constant tone extended CTE signal externally. It should be noted that the location coordinates of each transmitter are pre-set and known. Furthermore, the constant tone extended CTE signal is obtained by appending a CTE segment to the end of the signal transmitted by the Bluetooth 5.1-based Bluetooth tag. The appended CTE segment is a continuous carrier signal with a constant frequency and constant modulation. This resulting constant tone extended signal contains multiple sampling points, which can support high-frequency sampling and IQ data acquisition at the receiver.
[0035] It should be noted that the IQ data sequence contains IQ samples collected at multiple sampling points, and each IQ sample distribution includes an I component and a Q component; wherein the I component is the in-phase component; the Q component is the orthogonal component.
[0036] Furthermore, the receiving end is a receiving device that supports Bluetooth 5.1 and is equipped with a multi-antenna array, which is used to capture the CTE signal transmitted by the transmitter and extract the IQ data sequence therefrom.
[0037] It should be noted that the IQ data sequence can be expressed as:
[0038] z=(i1,q1,i2,q2,…,i n ,q n );
[0039] Where z is the IQ data sequence; n is the logarithm of the IQ data sequence; i k Expand the I component of the CTE signal at the kth sampling point for each constant tone; q k Expand the Q component of the k-th sampling point of the CTE signal for each constant tone.
[0040] Optionally, in a possible implementation, the specific step of obtaining multiple sets of positioning-related data includes:
[0041] Step 1: Generate a reference point location set based on the pre-set reference point layout strategy.
[0042] In this step, the pre-set reference point distribution strategy is set according to the actual structure of the ship's cabin, the complexity of the layout, and the desired positioning accuracy. It can be a regular grid distribution strategy or a non-uniform dense distribution strategy. It should be noted that in one embodiment, regular grid reference points can be arranged at fixed intervals in the cabin to form a uniform regular grid distribution strategy, which is suitable for open or regular cabin areas; in another embodiment, high-density reference points can be arranged in key areas of the ship, and low-density reference points can be arranged in ordinary areas of the ship to form a non-uniform dense distribution strategy to optimize resource allocation and adapt to the complex layout of the ship. For example, in one possible implementation, the spacing between reference points can be set to be less than or equal to 0.5m to form a uniform regular grid distribution strategy.
[0043] At the same time, for any reference point set in the layout strategy, record the position coordinates (x k ,y k ) as a unique identifier to obtain the set of all preset reference point positions.
[0044] Step 2: For each reference point position in the reference point position set, a transmitter is arranged at the reference point, and a receiving end is arranged at a specific position on the hull.
[0045] In this step, for each reference point in the reference point position set obtained in step 1, a transmitter is arranged at the position of the reference point to transmit a constant tone extended CTE signal, and a receiving end is arranged at a fixed position on the hull to capture the CTE signal and extract the IQ data sequence therein.
[0046] Step 3: For each of the preset multiple sampling scenarios, perform multiple rounds of sampling on each reference point position, and in each round of sampling, collect the IQ data sequence obtained by the receiving end in the process of receiving the CTE signal from the transmitter; wherein the preset multiple sampling scenarios are different, and the preset multiple sampling scenarios cover scenarios in which any of the following situations exist, and scenarios in which none of the following situations exist: personnel movement, equipment startup, and electromagnetic interference; the multiple rounds of sampling cover different time periods, and the hull is in different postures during the different time periods.
[0047] In this step, multiple sampling scenarios are pre-set for various scenarios during the ship's voyage. Furthermore, for each of the multiple sampling scenarios, multiple rounds of sampling are performed at each reference point, covering multiple time periods during the ship's operation to capture the various ship attitudes (possible ship attitudes include tilt, roll, and other attitude changes caused by waves, engine vibration, etc.) during these different time periods. The IQ data sequence acquired by the receiver during the process of receiving the CTE signal from the transmitter is collected. For example, in one embodiment, in the pre-set "device on + electromagnetic interference" sampling scenario, 10 rounds of sampling are performed at reference point P1 on the aft deck.
[0048] It should be noted that, based on the CTE signal collected in this step, a three-dimensional IQ data sequence covering dynamic scene interference, hull spatial position, and ship attitude changes can be extracted to simulate the fluctuation of the IQ data sequence during the actual operation of the ship and improve the robustness of the collected data.
[0049] Step 4: Bind each IQ data sequence collected at each reference point position to the reference point position to obtain a set of positioning related data.
[0050] In this step, for each reference point position, all IQ data sequences sent from the reference point position collected by the receiver are bound to the reference point position, thereby obtaining a set of positioning related data. For example, the complete IQ data sequence received by the receiver from a reference point device is z k =(i1,q1,i2,q2,…,i n ,q n ), the real space position label of the reference point device is Label k =(xk ,y k ), thus, a set of positioning related data (z k ,Label k ).
[0051] It should be noted that by repeating the above steps, multiple sets of positioning related data can be obtained.
[0052] Furthermore, in a possible implementation, after acquiring multiple sets of positioning-related data, the method further includes:
[0053] (1) Use statistical analysis methods to detect whether there are missing values in each set of positioning-related data, and fill in or eliminate the data with missing values.
[0054] In this step, each of the multiple sets of positioning-related data is checked for null values or incomplete records. For example, missing values in each set of positioning-related data can be identified through statistical analysis methods such as calculating the proportion of non-null values or visualizing the distribution of missing values. Furthermore, for positioning-related data detected to contain missing values, an appropriate imputation method, such as mean imputation, median imputation, interpolation, or model-based predictive imputation, is selected based on the missing value ratio and data characteristics. If the missing value ratio is too high or imputation may introduce significant deviations, the corresponding data records are removed.
[0055] (2) Using statistical threshold method or machine learning-based anomaly detection algorithm, it is identified that the IQ data sequence in each set of positioning-related data has abnormal jump characteristics, and the detected abnormal samples are eliminated or corrected.
[0056] In this step, after completing the missing value processing, it is necessary to further identify whether the IQ data sequence in each set of positioning-related data has abnormal jump characteristics through a statistical threshold method or a machine learning-based anomaly detection algorithm. Among them, the statistical threshold method usually sets the upper and lower limits of a specific indicator and marks data points that exceed the threshold as abnormal. For example, assuming that a set of IQ data sequences represents signal strength, with values of [50, 52, 51, 80, 53, 49, 100, 51], and the threshold range is [51, 101], then data point 50 is not within the threshold range and is marked as an abnormal data point. For another example, a machine learning-based anomaly detection algorithm identifies jump characteristics that deviate from the normal pattern by learning the normal distribution pattern of the data. For example, assuming that a set of IQ data sequences represents signal strength, with values of [50, 52, 51, 80, 53, 49, 100, 51], since most values are around 50, and 100 significantly deviates from the cluster area of other data points, 100 is identified as a jump characteristic that deviates from the normal pattern.
[0057] Furthermore, the detected abnormal samples can be eliminated to avoid interference with subsequent analysis, or corrected through smoothing, interpolation and other methods. In this way, the continuity of the positioning-related data can be retained, the quality of the positioning-related data can be improved, and the accuracy of subsequent analysis based on the positioning-related data can be improved.
[0058] The feature-screening-based positioning method provided in this embodiment for a ship environment employs a regular gridding or non-uniform dense placement strategy. Bluetooth tags are deployed at preset reference points to transmit constant-tone extended (CTE) signals. A receiver equipped with a multi-antenna array captures the signals and extracts IQ data sequences and angle-of-arrival information. Subsequently, through multi-scenario and multi-round sampling, encompassing dynamic scenarios such as personnel movement, device activation, electromagnetic interference, and varying ship postures, multiple sets of positioning-related data, including IQ data sequences and reference point locations, are acquired. Statistical analysis and anomaly detection methods are used to fill missing values and correct for abnormal jumps in the data. This effectively constructs a highly robust positioning dataset that can effectively simulate the complex environmental interference experienced in ship operation, improving the quality of the positioning dataset and its accuracy.
[0059] S102: Using the IQ data sequence in each set of positioning-related data as an input feature and the position in the set of positioning-related data as a label, construct a first training sample set using the multiple sets of positioning-related data.
[0060] In this step, it is assumed that the IQ data sequence s in the i-th group of positioning related data is i As input features, the set of reference points in the positioning data is located at the position t k As a label, the IQ data sequence in the i-th group of positioning related data can be expressed by the following formula:
[0061] s i =(i1,q1,i2,q2,…,i n ,q n );
[0062] Among them, s i ∈R 2n , is the IQ feature vector of the i-th sample.
[0063] Furthermore, the position t of the reference point in the i-th group of positioning related data i It can be expressed by the following formula:
[0064] t i ∈{1,2,…,C};
[0065] Among them, t i is the position of the reference point in the i-th group of positioning related data, 1, 2, ..., C are the position labels of C reference points respectively.
[0066] In this way, the first training sample set can be constructed using multiple sets of positioning-related data:
[0067]
[0068] Among them, D train is the first training sample set; m is the number of positioning related data.
[0069] S103: Train a gradient decision tree model using the first training sample set, and after the training is completed, determine an importance score for each feature in the input features based on the training results.
[0070] In this step, the first training sample set obtained in the above step is used to train the gradient decision tree model using the LightGradient Boosting Machine (LightGBM) algorithm to learn the gradient decision tree model from the input feature vector s. i To space tag t i The mapping relationship.
[0071] Specifically, the gradient boosting tree algorithm constructs a sequence of T decision trees {T1, T2, ..., T T ,}, the model is gradually optimized by gradient improvement; among them, the first lesson decision tree in the gradient decision tree model is a certain IQ feature vector s from the input i To the predicted position t i The tree fits the error of this round of prediction, and according to the error fitted by the tree, guides the prediction result of the decision tree at the next moment to gradually reduce the error and realize the training of the gradient decision tree model.
[0072] Furthermore, after the gradient decision tree model is trained, the importance score of each feature in the input features can be determined based on the training results according to the following formula:
[0073]
[0074] Among them, I j Score the importance of feature j; f j is the jth input feature; Gain(f j ,t) is the information gain brought by the j-th input feature when it is used as a split node in the t-th tree; T is the total number of decision trees in the gradient decision tree model.
[0075] It should be noted that the calculation of information gain can be based on loss functions such as the Gini index, cross entropy or square error, and is not limited in this application. Furthermore, the higher the importance score of an input feature, the greater the effect of the feature on improving the classification accuracy.
[0076] S104. For each feature, use the trained gradient decision tree model to evaluate the predictive ability of the feature under independent input conditions to obtain an overall stability score for the feature; wherein the overall stability score is used to characterize the stability of the feature in maintaining predictive accuracy at different positions.
[0077] It should be noted that the overall stability score measures the accuracy of the feature when it independently participates in positioning prediction at different positions. The higher the overall stability score of the feature, the more it can stably support accurate prediction under multi-position and multi-sample conditions, and has strong versatility and robustness. A low overall stability score of the feature indicates that the prediction effect based on the feature depends on a specific position, has poor stability, and has a limited scope of application.
[0078] Specifically, in one possible implementation, the use of a trained gradient decision tree model to evaluate the predictive ability of the feature under independent input conditions to obtain an overall stability score of the feature includes:
[0079] Step 1: For each feature, the values of other features except the feature in each training sample in the training sample set are set to 0 to generate a single-feature input data set corresponding to the feature.
[0080] In this step, for each feature to be evaluated, all samples in the training sample set are traversed, and all other feature values in the current sample except for the feature are forced to be zero, thereby generating a dedicated dataset that only retains the original data of the feature. For example, for a training sample, Among them, for s i =(i1,q1,i2,q2,…,i n ,q n ), the single-feature input data corresponding to this feature can be expressed as (i1,0,0,0,…,0,0).
[0081] Step 2: For each location of all transmitters included in the multiple sets of positioning-related data, use the trained gradient decision tree model to predict the single-feature input data corresponding to each feature to obtain the prediction result of the feature at that location.
[0082] Specifically, combining the above example, for example, for t i , find the single feature input data corresponding to all training samples at this position from all training samples, and input the single feature input data into the trained gradient decision tree model. For example, input (i1,0,0,0,…,0,0) into the gradient decision tree model to obtain the prediction result of the feature at this position.
[0083] Step 3: Determine the local stability score of the feature at the position based on the prediction result of the feature at the position; wherein the local stability score of the feature at the position is equal to the ratio of the number of samples effectively predicted by the feature at the position to the total number of samples corresponding to the feature at the position.
[0084] Specifically, if the prediction result of the feature at the position is consistent with its actual position, it is a valid prediction sample, otherwise it is an invalid prediction sample.
[0085] Specifically, the local stability score of the feature at this location can be calculated according to the following formula:
[0086]
[0087] Among them, S f,l is the local stability score of feature f at position l; V f,l is the number of samples of effective prediction of feature f at position l; N f,l is the total number of samples corresponding to feature f at position l.
[0088] Step 4: Determine the overall stability score of the feature based on the local stability scores of the feature at each location.
[0089] In a specific implementation, in one embodiment, a weighted value may be calculated based on the local stability score of the feature at each position and the weight predicted to be set for each position, and the weighted value may be determined as the feature stability score of the feature.
[0090] Optionally, in a possible implementation, the specific implementation process of this step may include:
[0091] 1. Determine the weight of each position;
[0092] Specifically, in a possible implementation method, the specific implementation process of this step may include: for each position, determining the weight of the position based on the number of samples corresponding to the position in the first training sample set and the total number of samples included in the first training sample set; or determining the weight of each position based on the number of positions included in the first training sample set; wherein the weight of each position is the same.
[0093] Specifically, in this implementation, the weight of the position can be determined based on the number of samples corresponding to the position:
[0094]
[0095] Among them, w l is the weight of position l; N lis the number of samples corresponding to the position in the first training sample set; N total is the total number of samples contained in the first training sample set.
[0096] In addition, if there are M transmitter positions, a fixed weight setting can be made so that the weight of each position is 1 / M.
[0097] It should be noted that the above two different weight determination methods are both related to the number of samples, and the two methods are applicable to different scenarios respectively; among them, the first weight determination method based on the number of samples is applicable to the scenario where the samples in the first training sample set are unevenly distributed; the second weight determination method of equal weight is applicable to the scenario where the samples in the first training sample set are evenly distributed.
[0098] Furthermore, in another possible implementation, determining the weight of each position further includes:
[0099] (1) Based on the importance level set by the user for each location and the correspondence between the preset importance level and the original weight, the original weight of each location is determined.
[0100] Specifically, the user sets an importance level for each location, for example, cabin = emergency level, deck = important level, rest area = ordinary level; further, the correspondence between the importance level and the original weight is pre-set, for example, in one embodiment, the preset weight of the emergency level is 0.7, the preset weight of the important level is 0.5, and the preset weight of the ordinary level is 0.3. In this way, the original weight of each location is obtained.
[0101] (2) Normalize the original weight of each position to obtain the weight of each position.
[0102] Specifically, the original weight of each position can be normalized according to the following formula to obtain the weight of each position:
[0103]
[0104] Among them, w l is the weight of position l; is the original weight of position l; is the sum of the original weights of all positions.
[0105] It should be noted that the specific method for determining the weight can be selected according to actual needs and is not limited in this application.
[0106] In this step, the original weights at each position are normalized into a standard, easy-to-understand and easy-to-use probability distribution format. This can significantly improve the interpretability of the weights and the convenience and mathematical consistency of downstream applications, while also reflecting the relative importance differences between the original weights.
[0107] 2. Based on the weight of each position, the local stability score of the feature at each position is weighted to obtain the overall stability score of the feature.
[0108] Specifically, in this step, the overall stability score of the feature can be obtained according to the following formula:
[0109]
[0110] in, is the overall stability score of feature f at position l; w l is the weight of position l; S f,l Score the local stability of feature f at location l.
[0111] S105 , determining a comprehensive score for each feature according to the importance score and the overall stability score of each feature, and screening a target feature set for positioning from multiple features included in the input features based on the comprehensive score of each feature.
[0112] In this step, before determining the comprehensive score of each feature, the overall stability score and importance score of the feature need to be normalized (for example, by performing normalization based on Min-Max normalization).
[0113] Specifically, the importance score and the overall stability score can be normalized using the Min-Max normalization process according to the following formula:
[0114]
[0115] in, is the normalized score of the jth feature, ranging from [0, 1]; is the minimum value of the score set of the j-th feature at each position; is the maximum value of the score set of the j-th feature at each position.
[0116] It should be noted that the above formula is a general formula for Min-Max normalization. When the importance score and the overall stability score are standardized respectively, each variable will also be adjusted accordingly. For example, according to the above formula, if the importance score of the j-th feature is standardized, the obtained importance score of the j-th feature after standardization is is the minimum value of the importance score set of the j-th feature at each position; is the maximum value of the importance score set of the j-th feature at each position; if the overall stability score of the j-th feature is standardized, the overall stability score of the j-th feature after standardization is is the minimum value of the overall stability score set of the j-th feature at each position; is the maximum value of the overall stability score set of the j-th feature at each position.
[0117] The comprehensive score of each feature is determined based on the importance score and stability score of each feature, including:
[0118]
[0119] Among them, R j is the comprehensive score of the jth feature; is the importance score of the jth feature after normalization; α is the weight coefficient of the importance score; is the overall stability score of the j-th feature after standardization; β is the weight coefficient of the overall stability score.
[0120] It should be noted that the above formula for determining the comprehensive score of each feature uses a weighted linear combination method to calculate the comprehensive score. This calculation method can dynamically adjust the weights of the importance score and stability score according to actual needs, which is convenient for interpretation and tuning.
[0121] Furthermore, in another embodiment, the comprehensive score of each feature may be determined according to the importance score and stability score of each feature according to the following formula:
[0122]
[0123] Among them, R j is the comprehensive score of the jth feature; is the importance score of the jth feature after normalization; α is the weight coefficient of the importance score; is the overall stability score of the j-th feature after standardization; β is the weight coefficient of the overall stability score.
[0124] It should be noted that using a weighted geometric mean to determine the comprehensive score for each feature is more suitable for scenarios requiring both importance and stability scores, and can provide greater environmental robustness. For example, in one embodiment, α = 0.7 and β = 0.3, where the focus is on positioning accuracy; in another embodiment, α = 0.4 and β = 0.6, where the focus is on the model's ability to adapt to a variety of scenarios.
[0125] Furthermore, based on the comprehensive score of each feature, a target feature set for positioning is screened out from the multiple features included in the input features.
[0126] Specifically, a comprehensive score threshold is preset. If the comprehensive score of a feature among the input features is higher than the comprehensive score threshold, it is selected and included in the target feature set. Conversely, if the comprehensive score of a feature is lower than the comprehensive score threshold, it is removed and not selected into the target feature set. For example, in one embodiment, if the comprehensive score threshold is 80, features with a comprehensive score higher than 80 are selected and included in the target feature set for positioning.
[0127] In addition, this step also sets the upper limit N of the number of features contained in the target feature set. max , in order to limit the feature dimension of the target feature set used for positioning and control the model complexity of the positioning model finally trained. For example, in one embodiment, N max is 20. In specific implementation, if the number of target features screened out according to the above method is greater than 20, the first 20 target features can be selected in descending order of comprehensive scores.
[0128] It should be noted that by evaluating each feature at each location, calculating the local stability score for each feature at different spatial locations, and further integrating it into an overall stability score, we can effectively characterize the consistency of the feature's performance in a changing environment, avoiding the problem of only considering the global average prediction performance while ignoring spatial differences. The selected features are more stable in prediction across scenarios, thereby improving the model's generalization performance in new locations or dynamic environments. Furthermore, it is understandable that some features may perform well in some locations (for example, due to the coincidence of specific multipath patterns resulting in locally high discrimination) but perform extremely poorly in other areas. If evaluated solely based on importance or global error, such "locally advantageous" features may be incorrectly selected. The method provided in this implementation evaluates the proportion of valid prediction samples at each location to eliminate such features that lack stability in the overall space, thereby improving the reliability of feature selection. The final target feature set is selected based on a dual evaluation of "importance score" and "stability score," effectively reducing the negative impact of redundant features and unstable inputs on model training. This not only improves the prediction accuracy of the final positioning model, but also enhances the model's adaptability to new samples, boundary areas, and channel fluctuations, meeting the dual requirements of accuracy and stability in actual deployment.
[0129] S106: Construct a second training sample set for training a positioning model based on the target feature set and the training sample set, and train the positioning model using the second training sample set.
[0130] In this step, based on the target feature set screened out for positioning, for the training sample set, other features in the training sample set except the input features in the target feature set are removed, and the input features contained in the target feature set are retained. For example, the features in the training sample set include 40 features, and the target feature set only contains 12 features. Therefore, the remaining 28 features in the training sample set except the 12 features in the target feature set are removed, and the retained features are used as high-value features to construct a second training sample set, and the second training sample set is used to train the positioning model. In this way, not only can the positioning error of the positioning model be effectively reduced, but also the redundant information caused by unnecessary low-value features can be avoided.
[0131] S107 . After obtaining the data to be predicted, extracting feature values corresponding to the target feature set from the data to be predicted, and inputting the extracted feature values into the positioning model so that the positioning model predicts the position corresponding to the data to be predicted.
[0132] Specifically, after receiving the Bluetooth IQ data to be predicted, the feature values corresponding to the target feature set are first dynamically extracted, and other unnecessary redundant features are eliminated; then the extracted optimized feature vectors are input into the pre-trained positioning model, and the positioning model predicts the three-dimensional spatial coordinates corresponding to the data to be predicted.
[0133] The positioning method based on feature screening in a ship environment provided in this embodiment eliminates irrelevant or unstable features by combining the feature importance score with the overall stability score, retains key features with high discrimination for the positioning task, and effectively improves the accuracy of the predicted position. In addition, when determining the overall stability score of the feature, the consistency of the performance of the feature in different positions is comprehensively considered, which can improve the generalization ability of the model in different cabins and different environments and reduce prediction fluctuations; further, by reducing the input dimension through feature screening and reducing the training and inference burden of the model, the real-time performance of the system can be improved, and it can adapt to the resource-constrained edge computing deployment requirements in the ship environment.
[0134] Corresponding to the aforementioned embodiment of a positioning method based on feature screening in a ship environment, the present application also provides an embodiment of a positioning device based on feature screening in a ship environment.
[0135] The embodiment of the feature-based positioning device in a ship environment of the present application can be applied to a feature-based positioning device in a ship environment. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of the feature-based positioning device in the ship environment in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory and running them. From the hardware level, if Figure 3 As shown, this is a hardware structure diagram of a positioning device based on feature screening in a ship environment of the present application, except Figure 3 In addition to the processor, memory, network interface, and non-volatile memory shown, the positioning device used in the ship environment in which the device in the embodiment is located may also include other hardware based on the actual function of the positioning device based on feature screening in the ship environment, which will not be described in detail.
[0136] Figure 2 This is a structural diagram of the first embodiment of the positioning device based on feature screening in a ship environment provided by this application. Figure 2 The device provided in this embodiment includes an acquisition module 210, a construction module 220, a determination module 230 and a prediction module 240; wherein,
[0137] The acquisition module 210 is configured to acquire multiple sets of positioning-related data; each set of positioning-related data includes an IQ data sequence acquired by the receiving end during the process of receiving a constant tone extension (CTE) signal transmitted from a transmitter, and the location of the transmitter;
[0138] The construction module 220 is configured to use the IQ data sequence in each set of positioning-related data as an input feature and the position in the set of positioning-related data as a label, and use the multiple sets of positioning-related data to construct a first training sample set;
[0139] The determination module 230 is configured to train a gradient decision tree model using the first training sample set, and after the training is completed, determine the importance score of each feature in the input features based on the training results;
[0140] The determination module 230 is configured to evaluate the predictive ability of each feature under independent input conditions using a trained gradient decision tree model to obtain an overall stability score for the feature; wherein the overall stability score is used to indicate the degree to which the feature maintains predictive accuracy at different positions;
[0141] The determination module 230 is configured to determine a comprehensive score for each feature based on the importance score and stability score of each feature, and to select a target feature set for positioning from the multiple features included in the input features based on the comprehensive score of each feature;
[0142] The construction module 220 is configured to construct a second training sample set for training a positioning model based on the target feature set and the training sample set, and train the positioning model using the second training sample set;
[0143] The prediction module 240 is used to extract feature values corresponding to the target feature set from the data to be predicted after obtaining the data to be predicted, and input the extracted feature values into the positioning model so that the positioning model predicts the position corresponding to the data to be predicted.
[0144] The device of this embodiment can be used to perform Figure 1 The steps, specific implementation principles and implementation processes of the method embodiment shown are similar and will not be repeated here.
[0145] Please continue to refer to Figure 3 The present application also provides a positioning device based on feature screening in a ship environment, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps of any one of the methods provided in the first aspect of the present application are implemented.
[0146] The present application also provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any one of the methods provided in the present application when the program is executed by a processor.
[0147] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0148] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present application scheme. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0149] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A positioning method based on feature screening, characterized in that: The positioning method is applied to a ship environment; the method comprises: Acquire multiple sets of positioning-related data; wherein each set of positioning-related data includes an IQ data sequence acquired by a receiving end in the process of receiving a constant tone extension (CTE) signal transmitted from a transmitter, and the location of the transmitter; Using the IQ data sequence in each set of positioning-related data as an input feature and the position in the set of positioning-related data as a label, and using the multiple sets of positioning-related data to construct a first training sample set; Training a gradient decision tree model using the first training sample set, and after the training is completed, determining an importance score for each feature in the input features based on the training results; For each feature, the trained gradient decision tree model is used to evaluate the predictive ability of the feature under independent input conditions to obtain the overall stability score of the feature; wherein the overall stability score is used to indicate the stability of the feature in maintaining predictive accuracy under different positions; Determining a comprehensive score for each feature according to the importance score and the stability score of each feature, and screening a target feature set for positioning from multiple features included in the input features based on the comprehensive score of each feature; Based on the target feature set and the training sample set, construct a second training sample set for training a positioning model, and train the positioning model with the second training sample set; After the data to be predicted is acquired, feature values corresponding to the target feature set are extracted from the data to be predicted, and the extracted feature values are input into the positioning model so that the positioning model predicts the position corresponding to the data to be predicted.
2. The method according to claim 1, characterized in that For each feature, the trained gradient decision tree model is used to evaluate the predictive ability of the feature under independent input conditions to obtain the overall stability score of the feature, including: For each feature, the values of other features except the feature in each training sample in the training sample set are set to 0 to generate a single-feature input data set corresponding to the feature; For each location of all transmitter locations included in the multiple sets of positioning-related data, using the trained gradient decision tree model, predict the single-feature input data corresponding to each feature to obtain a prediction result of the feature at the location; Determine the local stability score of the feature at the position based on the prediction result of the feature at the position; wherein the local stability score of the feature at the position is equal to the ratio of the number of samples effectively predicted by the feature at the position to the total number of samples corresponding to the feature at the position; The overall stability score of the feature is determined based on the local stability scores of the feature at each location.
3. The method according to claim 1, characterized in that The comprehensive score of each feature is determined based on the importance score and overall stability score of each feature, including: Among them, R j is the comprehensive score of the jth feature; is the importance score of the jth feature after normalization; α is the weight coefficient of the importance score; is the overall stability score of the j-th feature after standardization; β is the weight coefficient of the overall stability score.
4. The method according to claim 1, wherein Determining the overall stability score of the feature based on the local stability scores of the feature at each position includes: Determine the weight of each position; The local stability score of the feature at each position is weighted based on the weight of each position to obtain the overall stability score of the feature.
5. The method according to claim 4, characterized in that Determining the weight of each position includes: For each position, determine a weight for the position according to the number of samples corresponding to the position in the first training sample set and the total number of samples included in the first training sample set; or, The weight of each position is determined according to the number of positions included in the first training sample set; wherein the weight of each position is the same.
6. The method according to claim 4, characterized in that Determining the weight of each position includes: Determine the original weight of each location based on the importance level set by the user for each location and the correspondence between the preset importance level and the original weight; The original weight of each position is normalized to obtain the weight of each position.
7. The method according to claim 1, characterized in that The obtaining of multiple sets of positioning related data includes: Generate a reference point location set based on a pre-set reference point layout strategy; For each reference point position in the reference point position set, a transmitter is arranged at the reference point, and a receiver is arranged at a specific position on the hull; For each of the preset multiple sampling scenarios, multiple rounds of sampling are performed at each reference point position, and in each round of sampling, an IQ data sequence obtained by the receiving end during the process of receiving the CTE signal from the transmitter is collected; wherein the preset multiple sampling scenarios are different, and the preset multiple sampling scenarios cover scenarios in which any of the following situations exist, and scenarios in which none of the following situations exist: personnel movement, equipment startup, and electromagnetic interference; the multiple rounds of sampling cover different time periods, and the hull is in different postures during the different time periods; Each IQ data sequence collected at each reference point is bound to the reference point position to obtain a set of positioning related data.
8. The method according to claim 1 or 7, characterized in that After acquiring the multiple sets of positioning-related data, the method further includes: Use statistical analysis methods to detect whether there are missing values in each set of positioning-related data, and fill or eliminate the data with missing values; The statistical threshold method or the anomaly detection algorithm based on machine learning is used to identify whether the IQ data sequence in each set of positioning-related data has abnormal jump characteristics, and the detected abnormal samples are eliminated or corrected.
9. A positioning device based on feature screening in a ship environment, characterized in that: The device includes an acquisition module, a construction module, a determination module and a prediction module; wherein, The acquisition module is configured to acquire multiple sets of positioning-related data; wherein each set of positioning-related data includes an IQ data sequence acquired by the receiving end in the process of receiving a constant tone extension (CTE) signal transmitted from a transmitter, and the location of the transmitter; The construction module is configured to use the IQ data sequence in each set of positioning-related data as an input feature and the position in the set of positioning-related data as a label, and to construct a first training sample set using the multiple sets of positioning-related data; The determination module is configured to train a gradient decision tree model using the first training sample set, and after the training is completed, determine the importance score of each feature in the input features based on the training results; The determination module is configured to evaluate the predictive ability of each feature under independent input conditions using a trained gradient decision tree model to obtain an overall stability score for the feature; wherein the overall stability score is used to indicate the degree to which the feature maintains predictive accuracy at different positions; The determination module is configured to determine a comprehensive score for each feature based on the importance score and stability score of each feature, and to select a target feature set for positioning from a plurality of features included in the input features based on the comprehensive score of each feature; The construction module is configured to construct a second training sample set for training a positioning model based on the target feature set and the training sample set, and train the positioning model using the second training sample set; The prediction module is used to extract feature values corresponding to the target feature set from the data to be predicted after obtaining the data to be predicted, and input the extracted feature values into the positioning model so that the positioning model predicts the position corresponding to the data to be predicted.
10. A positioning device based on feature screening in a ship environment, characterized in that: The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of any one of the methods provided in the first aspect of the present application are implemented.