Ship position inference method based on Sentecnce-Bert and HMM model
By combining Sentence-BERT and HMM models, geolocation information is screened out from social media data, and the problem of ship position inference when AIS data is unreliable is solved, achieving more efficient position inference and reducing the workload of manual analysis.
Patent Information
- Application Number
- CN202411884997.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-06
AI Technical Summary
In marine vessel surveillance and navigation, it is difficult to accurately infer the position of the ship when AIS data is unreliable, especially when the ship tries to maintain concealment.
Using the method based on Sentence-BERT and HMM models, data that can reflect specific geographical location information is filtered out from massive unstructured social media data, text embedding vectors are generated through Sentence-BERT, and geographic location inference is used to use the HMM model.
It improves the efficiency of ship position inference when AIS signals are unavailable or unreliable, reduces the workload of manual screening of data, and provides a complement to conventional means of judgment.
Smart Images

Figure CN119938947A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine learning technology, and in particular to a method for inferring ship position based on Sentecnce-Bert and HMM models. Background Art
[0002] In modern maritime vessel monitoring and navigation, AIS (Automatic Identification System) data is an important data source for tracking and identifying ships. However, in some special cases, ships at sea need to remain concealed when conducting illegal operations. Turning off the AIS system or disguising as other ships to report AIS information is a common method to avoid exposing their location and action intentions, which brings certain difficulties to inferring their location.
[0003] In related technologies, social media information data has become a supplementary means of inferring the target's geographic location. Various social media software and websites have become important platforms for people to share their daily lives, opinions and information. Social media (hereinafter referred to as social media) data refers to various types of information collected from social media platforms, including user-posted content (posts, comments, pictures and videos, etc.), user behavior data (likes, shares, follows, etc.), user profiles, metadata (posting time, location, device information, etc.) and emotional data (positive, negative or neutral emotional tendencies). When working, shipboard staff sometimes share comments through social media platforms. These contents may contain information about the target's geographic location, which has a certain reference value for predicting the geographic location of the target of concern.
[0004] Therefore, screening out social media information that is relatively valuable for reference can greatly reduce the workload of manual analysis and can more conveniently and efficiently use social media information to infer the geographic location information of ship targets. Summary of the invention
[0005] In view of the above problems, the present application provides a ship position inference method based on Sentecnce-Bert and HMM models to achieve more accurate and efficient inference of the target geographical location.
[0006] In a first aspect, an embodiment of the present application provides a method for estimating a ship position based on a Sentecnce-Bert and HMM model, and the method for estimating a ship position based on a Sentecnce-Bert and HMM model comprises:
[0007] Acquiring social media data, wherein the social media data includes: social media data of relevant personnel corresponding to the target maritime vessel;
[0008] Input the social media data into a pre-trained Sentence-BERT model to generate text embedding vectors;
[0009] The text embedding vector is input into a pre-trained hidden Markov model to infer the geographic location of the target maritime vessel.
[0010] In some embodiments, the ship position inference method based on Sentecnce-Bert and HMM model further includes:
[0011] Preprocessing the social media data to obtain a standardized data set;
[0012] The standardized data set is annotated to generate an annotated data set containing annotation information.
[0013] In some embodiments, the pre-trained Sentence-BERT model comprises the following steps to determine:
[0014] Build the Sentence-BERT model;
[0015] The Sentence-BERT model is trained using a labeled dataset to obtain a pre-trained Sentence-BERT model and text embedding vector.
[0016] In some embodiments, the pre-trained hidden Markov model comprises the following steps to determine:
[0017] Constructing a vector annotation dataset based on the text embedding vector and annotation information;
[0018] The hidden Markov model is trained using a vector annotated data set to obtain a pre-trained hidden Markov model.
[0019] In some embodiments, annotating the standardized data set to generate an annotated data set containing annotation information includes:
[0020] Establishing a geographic location and tag comparison dictionary table based on the location information in the social media data;
[0021] The standardized data set is annotated using the geographic location and the tag reference dictionary table to generate an annotated data set containing annotated information.
[0022] In some embodiments, preprocessing the social media data to obtain a standardized data set includes:
[0023] Performing data cleaning on the social media data to obtain target social media data, wherein the data cleaning includes at least one of removing special characters, removing duplicate data, and removing noise data;
[0024] The target social media data is processed in a standard data format to obtain a standardized data set, wherein the standard data format includes at least one of a text ID, original text content, a publishing account, an associated target, a publishing time, and an indicated location.
[0025] In some embodiments, inputting the text embedding vector into a pre-trained hidden Markov model to infer the geographic location of the target marine vessel comprises:
[0026] Obtaining an output probability value obtained by inputting the text embedding vector into a pre-trained hidden Markov model;
[0027] When it is determined that the output probability value is greater than a threshold, the corresponding social media data is determined as reference data for inferring the geographic location of the target marine vessel.
[0028] In a second aspect, an embodiment of the present application provides a ship position inference device based on Sentecnce-Bert and HMM models, comprising:
[0029] An acquisition module is used to acquire social media data, wherein the social media data includes: social media data of relevant personnel corresponding to the target marine vessel;
[0030] A generation module, configured to input the social media data into a pre-trained Sentence-BERT model to generate a text embedding vector;
[0031] The inference module is used to input the text embedding vector into a pre-trained hidden Markov model to infer the geographic location of the target marine vessel.
[0032] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores a program code that can be executed on the processor, and when the program code is executed by the processor, a ship position inference method based on Sentecnce-Bert and HMM model as described in any implementation method of the first aspect is implemented.
[0033] In a fourth aspect, an embodiment of the present application provides a computer storage medium storing one or more programs, which can be executed by an electronic device as described in the third aspect to implement a ship position inference method based on Sentecnce-Bert and HMM model as described in any embodiment of the first aspect.
[0034] The embodiments of the present application provide a method, device, electronic device and storage medium for inferring the position of a ship based on the Sentence-BERT and HMM models. The method obtains social media data, wherein the social media data includes social media data of relevant personnel corresponding to the target maritime vessel, inputs the social media data into a pre-trained Sentence-BERT model to generate a text embedding vector, and inputs the text embedding vector into a pre-trained hidden Markov model to infer the geographic location of the target maritime vessel. When the AIS signal is unavailable or unreliable, the hidden Markov model and the Sentence-BERT model are used to filter out data that can reflect specific geographic location information from massive unstructured social media data. The method aims to improve the efficiency of inferring the geographic location of the maritime vessel target and reduce the workload of manual data screening. The method can be used as a supplement to conventional means of determining the target geographic location (satellite, radar, etc.).
[0035] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Hereinafter, the present application will be described in more detail based on embodiments and with reference to the accompanying drawings.
[0037] Figure 1 A schematic diagram of a method for estimating a ship's position based on Sentecnce-Bert and HMM models proposed in an embodiment of the present application is shown;
[0038] Figure 2 A schematic diagram of an inference process of an exemplary ship position inference model based on social media data proposed in one embodiment of the present application is shown;
[0039] Figure 3 A schematic diagram showing the composition of social media data fields in an exemplary standard data format proposed in an embodiment of the present application is shown;
[0040] Figure 4 The structure block diagram of a ship position estimation device based on Sentecnce-Bert and HMM model proposed in one embodiment of the present application is shown;
[0041] Figure 5 A structural block diagram of an electronic device for executing a method for estimating a ship position based on a Sentecnce-Bert and HMM model according to an embodiment of the present application is shown;
[0042] Figure 6A computer-readable storage medium for storing or carrying a method for inferring a ship position based on Sentecnce-Bert and an HMM model according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0043] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with embodiments and drawings. The exemplary embodiments of the present invention and their description are only used to explain the present invention and are not intended to limit the present invention.
[0044] Given the massive amount of social media data, manually screening useful information is a huge workload. By using automated analysis technology to screen out social media information that is relatively valuable for reference, the workload of manual analysis can be greatly reduced, and social media information can be used more conveniently and efficiently to infer the geographic location information of ship targets.
[0045] Research has found that in recent years, the development of natural language processing (NLP) technology, especially the emergence of pre-trained models based on deep learning (such as BERT and its variants), has provided new possibilities for understanding and utilizing complex semantic information in natural language. In particular, the Sentence-BERT (SBERT) model can efficiently generate semantically rich embedding vectors through the improvement of BERT (Bidirectional Encoder Representations from Transformers), providing a new tool for in-depth understanding of text content. As a powerful statistical model, the Hidden Markov Model (HMM) is widely used in the analysis of time series data, and is particularly good at processing and inferring the hidden state behind the observed data.
[0046] In view of the above technical problems, the inventors have analyzed and combined the efficient semantic representation ability of SBERT and the advantages of HMM in time series analysis, and proposed a ship position inference method, device, electronic device and storage medium based on Sentecnce-Bert and HMM model, which can realize the inference of the target geographical location. The present invention aims to solve the limitations and challenges in the prior art through this innovative method.
[0047] The ship position inference method, device, electronic device and storage medium based on Sentecnce-Bert and HMM models provided in the embodiments of the present application can use hidden Markov models and Sentence-BERT models to filter out data that can reflect specific geographic location information from massive unstructured social media data when AIS signals are unavailable or unreliable, aiming to improve the efficiency of inferring the geographic location of maritime ship targets and reduce the workload of manual data screening. This method can be used as a supplement to conventional means of judging the target geographic location (satellite, radar, etc.). Among them, the ship position inference method based on Sentecnce-Bert and HMM models is described in detail in subsequent embodiments.
[0048] The following is an introduction to the application scenarios of the ship position inference method based on Sentecnce-Bert and HMM model provided in the embodiment of the present application:
[0049] See also Figure 1 , Figure 1 FIG. 1 is a flow chart of a method for estimating a ship position based on a Sentecnce-Bert and HMM model provided in an embodiment of the present application. In this embodiment, the method for estimating a ship position based on a Sentecnce-Bert and HMM model can be applied to the following: Figure 4 The ship position estimation device 300 based on Sentecnce-Bert and HMM model is shown in FIG. Figure 5 In the electronic device 200 shown, the electronic device may include one or more electronic devices, and information may be transmitted between the multiple electronic devices by wireless and / or wired means. The multiple electronic devices may collaborate to complete the ship position inference method based on Sentecnce-Bert and HMM models. For example, the electronic device may include a computer, a mobile terminal, a tablet, etc., which is not limited in this application. Figure 1 The process shown is described in detail, and the ship position inference method based on Sentecnce-Bert and HMM model may include S110 to S130.
[0050] S110: Acquire social media data, wherein the social media data includes: social media data of relevant personnel corresponding to the target maritime vessel.
[0051] In an embodiment of the present application, in the process of acquiring social media data: social media information posted by persons related to the observation target is collected. In some aspects, the observation target may be an element in the observation target set determined according to actual conditions. In other aspects, persons related to the observation target may be all elements of the target-associated person set determined according to actual conditions. In still other aspects, the collected social media data content may cover information such as text content, publishing account, and publishing time.
[0052] S120: Input social media data into the pre-trained Sentence-BERT model to generate text embedding vectors.
[0053] In an embodiment of the present application, in the present application, the social media information involved in the global location tracking technology is processed using the Sentence-BERT model to obtain a text embedding vector, so that the embedding vector that can reflect a specific target in a specific geographic location is spatially closer.
[0054] It should be noted that in order to achieve the above purpose, the model is fine-tuned and then trained using social media information data.
[0055] For example, the trained Sentence-BERT model is used to extract the semantic features of social media text data and generate text embedding vectors of social media information. The Sentence-BERT model takes the original content text of social media data as input and generates a text embedding vector of size 1×768 after model calculation.
[0056] S130: Input the text embedding vector into a pre-trained hidden Markov model to infer the geographic location of the target maritime vessel.
[0057] In the embodiment of the present application, the generated embedding vector is used as the input of the trained HMM model to infer the geographic location. Based on the input feature embedding vector and the geographic location information specified in the dictionary table, the HMM can output the probability of a specific observation target in a specified geographic area at a specified time.
[0058] In this embodiment, the Hidden Markov Model and the Sentence-BERT model are used to filter data reflecting specific geographic location information from massive unstructured social media data, thereby reducing the workload of manual data screening and improving efficiency.
[0059] In some embodiments, the ship position inference method based on Sentecnce-Bert and HMM models also includes S111 to S112.
[0060] S111: Preprocess the social media data to obtain a standardized data set.
[0061] In an embodiment of the present application, data preprocessing may include processes such as text cleaning of data and data standardization.
[0062] In some embodiments, S111 also includes S1111 to S1112.
[0063] S1111: Perform data cleaning on the social media data to obtain target social media data, wherein the data cleaning includes at least one of removing special characters, removing duplicate data, and removing noise data.
[0064] In the embodiment of the present application, data cleaning mainly involves removing special characters, duplicate data and noise data from the collected social media data. Special characters are characters that cannot be processed by the model, duplicate data is determined based on the repetitiveness of the text content, and noise data is data with more garbled characters and affects semantic understanding.
[0065] S1112: Process the target social media data in a standard data format to obtain a standardized data set, wherein the standard data format includes at least one of a text ID, original content, a publishing account, an associated target, a publishing time, and an indicated location.
[0066] In the embodiment of the present application, social media data standardization processing refers to converting the collected raw social media data into data in a standard field format according to the defined fields. The annotated social media information includes text data and metadata related to the text. For specific field definitions, see Tables 1 and Figure 3 , Table 1 is the standard format social media data field definition table, Figure 3 A schematic diagram of the composition of social media data fields in an exemplary standard data format proposed in this application.
[0067]
[0068] Table 1
[0069] It should be noted that the original content may include the original content of social media data posted by users, which is the main content of social media information, including posts, comments or messages posted by users. The language of the text information includes English, and the length is between 0 and 2048 bytes.
[0070] The posting account may include information that usually includes the social media account that posted the information, such as username, user ID or other identifier. The account user should be related to the key strategic target (service personnel, staff, etc.), and the target can be associated with the user's identity information.
[0071] The associated target may include recording the identity information of the key strategic target contained in the social media information, indicating that the text is related to a specific strategic target.
[0072] The release time may include a record that generally includes the date and time at which the information is released. In some implementations, the release time may be used to track the time series and trends of events.
[0073] The indicated location may include geographic information related to the port or sea area contained in the text. In some embodiments, the location information can be used to associate the determined activity location of the key strategic target for model annotation and training.
[0074] It should be noted that in addition to the preprocessing of social media information, it is also necessary to establish a dictionary table of targets and geographic locations, which can be used to define the boundaries of targets and geographic locations during the implementation of the technology.
[0075] S112: Annotate the standardized data set to generate an annotated data set containing annotation information.
[0076] In some embodiments, S112 also includes S1121 to S1122.
[0077] S1121: Establishing a geographic location and tag comparison dictionary table based on the location information in the social media data.
[0078] S1122: Annotate the standardized data set using the geographic location and the label reference dictionary table to generate an annotated data set containing annotation information.
[0079] In an embodiment of the present application, a geographic location and label comparison dictionary table is established based on the location information in the social media data, and the trained social media data are annotated one by one to generate an annotated data set containing the annotated information.
[0080] In some embodiments, the pre-trained Sentence-BERT model includes the following steps S113 to S114 to determine
[0081] S113: Build Sentence-BERT model.
[0082] S114: Use the labeled dataset to train the Sentence-BERT model to obtain a pre-trained Sentence-BERT model and text embedding vector.
[0083] In this embodiment, Sentence-BERT is based on the BERT model and is specifically used to process text semantic similarity and text embedding tasks. Sentence-BERT can map text sentences into vector representations with semantic information, making similar sentences closer in the vector space.
[0084] In some embodiments, the pre-trained hidden Markov model includes the following steps S115 to S116 to determine:
[0085] S115: Construct a vector annotation dataset based on the text embedding vector and annotation information.
[0086] S116: Use the vector annotated dataset to train the hidden Markov model to obtain a pre-trained hidden Markov model.
[0087] In the embodiment of the present application, based on the open source Sentence-BERT, the labeled social media data is used for model training. Afterwards, the text embedding vector generated by the Sentence-BERT model and the original annotation information are used to construct the vector annotation data used for HMM model training. Based on the open source HMM model, the vector annotation data is used for training, and finally a reasonable HMM model is generated.
[0088] In some embodiments, for the above-mentioned pre-trained Sentence-BERT model and the pre-trained hidden Markov model, the training of the two models is terminated when the training loss is lower than a specified threshold, and can pass the test of a specified test data set, and the test is passed when the accuracy reaches a specified value.
[0089] Among them, the test data set and training data set are subsets of the original labeled data set. The data set is divided into training data set and test data set according to a certain ratio. The model trained and tested by setting standards is used as a standard model for inferring unknown social media data.
[0090] In some embodiments, S130 also includes S131 to S132.
[0091] S131: Obtain an output probability value obtained by inputting the text embedding vector into a pre-trained hidden Markov model.
[0092] S132: When it is determined that the output probability value is greater than the threshold, the corresponding social media data is determined as reference data for inferring the geographic location of the target marine vessel.
[0093] In the embodiment of the present application, whether the corresponding area can be used as reference data is determined by calculating the output probability for the area.
[0094] Exemplarily, the HMM corresponding to the trained region R is used to output a probability value. If it is higher than 0.6, the social media data is considered to be event 1 data, where different numbers can correspond to different regions.
[0095] See also Figure 2 , Figure 2 A schematic diagram of the inference process of an exemplary ship position inference model based on social media data provided in an embodiment of the present application.
[0096] This is described by taking an exemplary implementation scenario as an example.
[0097] Exemplarily, the embodiment scenario includes 1 key monitoring target T, 1 activity area R, and two events (event 1: social media information can reflect that the target T is active in area R, event 0: social media information cannot reflect that the target T is active in area R).
[0098] The embodiment scenario includes 200 pieces of social media information, of which 80 pieces of social media information are event 1 data and 120 pieces of social media information are event 0 data.
[0099] The specific implementation steps S1-S7 are as follows:
[0100] S1: Data Collection
[0101] Perform data collection operations to obtain raw data from social media platforms.
[0102] S2: Data cleaning
[0103] Perform data cleaning operations on the text of social media data to remove special characters, duplicate data and noise data.
[0104] S3: Data Standardization
[0105] The original social media data is converted into standardized social media data according to the standard data format. The standard format data contains the following fields: text ID, original content, publishing account, associated target, publishing time, and indicated location. The field descriptions are shown in the following table.
[0106] S4: Data Annotation
[0107] According to the "indicated location" field in the annotated data and the constructed annotation information dictionary table, the annotated social media data is annotated to form annotated social media data. The annotation information dictionary table is as shown in the following table. Table 2 is the indicated location label correspondence table.
[0108] Indicate location Label Location R 1 unknown 0
[0109] Table 2
[0110] S5: Sentence-Bert model training
[0111] Select a suitable Sentence-Bert model from the public pre-trained model library, and use the labeled social media data for training to obtain the trained Sentence-Bert model.
[0112] S6: HMM model training
[0113] Use the trained Sentence-Bert model and the social media data in the above steps as input to generate text embedding vectors. Generate corresponding annotated text vectors based on the original labels of the social media data. Use the annotated text vectors to train the open source HMM model to obtain a trained HMM model.
[0114] S7: Text embedding vector generation
[0115] Load the trained Sentence-Bert model configuration and parameters and initialize the model. Use the loaded model to generate text embedding vectors for the selected 200 social media data one by one.
[0116] S8: Geographic Location Inference
[0117] Take the text embedding vector as input, use the HMM corresponding to the trained region R, and output the probability value. If it is higher than 0.6, it is considered that this social media data is event 1 data and can be used as reference data for human judgment.
[0118] In this application, by combining the efficient semantic representation capability of Sentence-BERT and the advantages of HMM in time series analysis, the geographical location of the target can be better inferred, the efficiency of inferring the geographical location of maritime ship targets can be improved, and the workload of manual data screening can be reduced.
[0119] See also Figure 4 , Figure 4 The structural block diagram of a ship position inference device based on Sentecnce-Bert and HMM model provided in the present application, the ship position inference device 300 based on Sentecnce-Bert and HMM model includes: an acquisition module 310, a generation module 320 and an inference module 330, wherein:
[0120] The acquisition module 310 is used to acquire social media data, wherein the social media data includes: social media data of relevant personnel corresponding to the target marine vessel.
[0121] The generation module 320 is used to input social media data into a pre-trained Sentence-BERT model to generate a text embedding vector.
[0122] The inference module 330 is used to input the text embedding vector into the pre-trained hidden Markov model to infer the geographic location of the target marine vessel.
[0123] The device embodiment in the present application may also include other modules, which specifically correspond to part of the content of the above method.
[0124] It should be noted that the device embodiments in the present application correspond to each other with the aforementioned method embodiments. The specific principles in the device embodiments can be found in the contents of the aforementioned method embodiments and will not be repeated here.
[0125] In several embodiments provided in this embodiment, the coupling between modules may be electrical, mechanical or other forms of coupling.
[0126] In addition, each functional module in each embodiment of the present invention may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of software functional modules.
[0127] See also Figure 5 , Figure 5 A structural block diagram of an electronic device 200 provided in an embodiment of the present application that can execute the above-mentioned ship position inference method based on Sentecnce-Bert and HMM model, the electronic device 200 can be a smart phone, tablet computer, computer or portable computer and other devices.
[0128] The electronic device 200 further includes a processor 202 and a memory 204. The memory 204 stores a program that can execute the contents of the aforementioned embodiments, and the processor 202 can execute the program stored in the memory 204.
[0129] Among them, the processor 202 may include one or more cores and message matrix units for processing data. The processor 202 uses various interfaces and lines to connect various parts of the entire electronic device 200, and executes various functions and processes data of the electronic device 200 by running or executing instructions, programs, code sets or instruction sets stored in the memory 204, and calling data stored in the memory 204. Optionally, the processor 202 can be implemented in at least one hardware form of digital signal processing (Digital Signal Processing, DSP), field programmable gate array (Field-Programmable Gate Array, FPGA), and programmable logic array (Programmable Logic Array, PLA). The processor 202 can integrate one or a combination of a central processing unit (Central Processing Unit, CPU), a graphics processing unit (Graphics Processing Unit, GPU) and a modulation decoder. Among them, the CPU mainly processes the operating system, user interface and application programs; the GPU is responsible for rendering and drawing display content; and the modem is used to process wireless communications. It can be understood that the above-mentioned modulation decoder may not be integrated into the processor, but may be implemented separately through a communication chip.
[0130] The memory 204 may include a random access memory (RAM) or a read-only memory (ROM). The memory 204 may be used to store instructions, programs, codes, code sets or instruction sets. The memory 204 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (e.g., instructions for a user to obtain a random number), instructions for implementing the following various method embodiments, etc. The data storage area may also store data (e.g., random numbers) created by the terminal during use, etc.
[0131] The electronic device 200 may also include a network module and a screen. The network module is used to receive and send electromagnetic waves, realize the mutual conversion between electromagnetic waves and electrical signals, and thus communicate with a communication network or other devices, such as communicating with an audio playback device. The network module may include various existing circuit components for performing these functions, such as antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, user identity modules (SIM) cards, memories, etc. The network module may communicate with various networks such as the Internet, corporate intranets, wireless networks, or communicate with other devices via wireless networks. The above-mentioned wireless networks may include cellular telephone networks, wireless local area networks, or metropolitan area networks. The screen may display interface content and perform data interaction.
[0132] Please refer to Figure 6 , Figure 6 The computer-readable storage medium 400 stores program code 410, which can be called by a processor to execute the method described in the above method embodiment.
[0133] The computer-readable storage medium 400 may be an electronic memory such as a flash memory, an EEPROM (electrically erasable programmable read-only memory), an EPROM, a hard disk, or a ROM. Optionally, the computer-readable storage medium includes a non-transitory computer-readable storage medium. The computer-readable storage medium 400 has storage space for program codes 410 that perform any method steps in the above method. These program codes 410 can be read from or written to one or more computer program products. The program codes 410 can be compressed, for example, in an appropriate form.
[0134] The embodiment of the present application also provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the ship position inference method based on Sentecnce-Bert and HMM model described in the above various optional implementations.
[0135] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features can be replaced by equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A ship position inference method based on Sentecnce-Bert and HMM model, characterized in that: The method comprises: Acquiring social media data, wherein the social media data includes: social media data of relevant personnel corresponding to the target maritime vessel; Input the social media data into a pre-trained Sentence-BERT model to generate text embedding vectors; The text embedding vector is input into a pre-trained hidden Markov model to infer the geographic location of the target maritime vessel.
2. The method for inferring ship position based on Sentecnce-Bert and HMM model according to claim 1, characterized in that: The ship position inference method based on Sentecnce-Bert and HMM model also includes: Preprocessing the social media data to obtain a standardized data set; The standardized data set is annotated to generate an annotated data set containing annotation information.
3. The method for inferring ship position based on Sentecnce-Bert and HMM model according to claim 2, characterized in that: The pre-trained Sentence-BERT model comprises the following steps: Build the Sentence-BERT model; The Sentence-BERT model is trained using a labeled dataset to obtain a pre-trained Sentence-BERT model and text embedding vector.
4. The method for inferring ship position based on Sentecnce-Bert and HMM model according to claim 3, characterized in that: The pre-trained hidden Markov model comprises the following steps: Constructing a vector annotation dataset based on the text embedding vector and annotation information; The hidden Markov model is trained using a vector annotated data set to obtain a pre-trained hidden Markov model.
5. The method for inferring ship position based on Sentecnce-Bert and HMM model according to claim 2, characterized in that: The step of labeling the standardized data set to generate a labeled data set containing labeling information includes: Establishing a geographic location and tag comparison dictionary table based on the location information in the social media data; The standardized data set is annotated using the geographic location and the tag reference dictionary table to generate an annotated data set containing annotated information.
6. The method for inferring ship position based on Sentecnce-Bert and HMM model according to claim 2, characterized in that: The preprocessing of the social media data to obtain a standardized data set includes: Performing data cleaning on the social media data to obtain target social media data, wherein the data cleaning includes at least one of removing special characters, removing duplicate data, and removing noise data; The target social media data is processed in a standard data format to obtain a standardized data set, wherein the standard data format includes at least one of a text ID, original text content, a publishing account, an associated target, a publishing time, and an indicated location.
7. The method for inferring ship position based on Sentecnce-Bert and HMM model according to claim 2, characterized in that: The step of inputting the text embedding vector into a pre-trained hidden Markov model to infer the geographic location of the target marine vessel comprises: Obtaining an output probability value obtained by inputting the text embedding vector into a pre-trained hidden Markov model; When it is determined that the output probability value is greater than a threshold, the corresponding social media data is determined as reference data for inferring the geographic location of the target marine vessel.
8. A ship position estimation device based on Sentecnce-Bert and HMM model, characterized in that: The device comprises: An acquisition module is used to acquire social media data, wherein the social media data includes: social media data of relevant personnel corresponding to the target marine vessel; A generation module, configured to input the social media data into a pre-trained Sentence-BERT model to generate a text embedding vector; The inference module is used to input the text embedding vector into a pre-trained hidden Markov model to infer the geographic location of the target marine vessel.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a program code that can be run on the processor, and when the program code is executed by the processor, the ship position inference method based on Sentecnce-Bert and HMM model as described in any one of claims 1-7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, and the program codes can be called by one or more processors to execute the ship position inference method based on Sentecnce-Bert and HMM model as described in any one of claims 1-7.