Maritime safety information lossless compression method and system based on context semantic coding
Through a method based on context semantic coding, maritime safety information is preprocessed, word frequency analysis and feature extraction, redundant information is identified and removed, and context information is used for dynamic probability prediction and adaptive encoding, which solves the problem of low compression efficiency of maritime safety information in the prior art, and realizes efficient and lossless compression transmission.
Patent Information
- Application Number
- CN202411840493.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-05-06
AI Technical Summary
The existing text compression algorithm has limited effect when processing small-scale and fixed-structure maritime safety information, and traditional methods have limitations when processing complex and variable text structures, which cannot meet the efficient transmission needs of maritime safety information.
The context-based semantic coding method is adopted to realize lossless compression by preprocessing, word frequency analysis, feature extraction and semantic analysis of the original maritime safety information, identifying and removing redundant information, and using context information for dynamic probability prediction and adaptive coding.
It improves the compression rate and transmission efficiency of maritime safety information, ensures that the compressed data can be completely restored to its original state, and meets the real-time and efficient transmission needs of maritime safety information.
Smart Images

Figure CN119945456A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information communication technology, and in particular to a method and system for lossless compression of maritime safety information based on context semantic coding. Background Art
[0002] The BeiDou Navigation Satellite System, as a global satellite navigation system independently built and operated by my country, not only provides core services such as navigation positioning, user detection, and standard timing, but also integrates short message communication functions. This function realizes data transmission through intersatellite links, ensuring all-weather and blind-spot coverage, which is particularly suitable for key areas such as life rescue and ocean monitoring. However, the transmission format of BeiDou short messages is fixed, and the length of a single message is limited, resulting in limited information transmission efficiency under limited bandwidth conditions.
[0003] In order to solve the bandwidth limitation problem in Beidou short message transmission, researchers have proposed a variety of data compression methods. These methods aim to remove redundant information and achieve efficient encoding so that information can be fully transmitted within a limited byte space. In the field of text compression, compression algorithms that have made significant progress mainly include dictionary coding, entropy coding, and predictive matching algorithms. For example, traditional dictionary coding methods such as LZ77 and LZMA, as well as entropy coding methods such as Huffman coding and arithmetic coding, have performed well in big data processing.
[0004] Although traditional compression algorithms are effective in processing big data, they are limited in their effectiveness when processing small-scale, fixed-structure maritime safety information, and may even increase space usage. In addition, although the Prediction by Partial Matching (PPM) series of algorithms perform well in compressing short texts and effectively reduce the redundant byte space of small text data through context modeling and frequency prediction, they still have limitations when processing complex and variable text structures. Although the PPMd algorithm further optimizes the flexibility of context matching based on PPM, it still has certain limitations. On the other hand, although the compression method based on neural networks can dynamically learn data distribution and achieve efficient compression, its storage and computational overhead is large and is not suitable for the real-time compression and transmission requirements of maritime information. Summary of the invention
[0005] The present invention proposes a method and system for lossless compression of maritime safety information based on context semantic coding, which solves the problem that the existing text compression algorithm is not suitable for maritime safety information with small data volume and high transmission speed requirements.
[0006] In order to solve the above technical problems, the present invention provides a method for lossless compression of maritime safety information based on context semantic coding, comprising the following steps:
[0007] Step S1: pre-processing the original maritime safety information, encoding the pre-processed maritime safety information according to the short message interface control file of the Beidou satellite area and the data format of the maritime safety information, and generating a maritime safety information message;
[0008] Step S2: performing a word frequency analysis on the pre-processed maritime safety information, constructing a maritime safety information vocabulary according to the word frequency analysis result, using the vocabulary to train a prediction model, using the trained prediction model to perform feature extraction and semantic analysis on the maritime safety information message, identifying and removing redundant information in the maritime safety information message, and obtaining a compact maritime safety information message;
[0009] Step S3: Using context information, dynamically probabilistically predict and adaptively encode characters in the compact maritime safety information message to achieve lossless compression.
[0010] Preferably, the preprocessing of the original maritime safety information in step S1 includes: performing text cleaning and grammar checking on the original maritime safety information, converting the cleaned maritime safety information into a standard format of maritime safety information content, and extracting data of key components of the original maritime information from the format-converted maritime safety information based on a trained neural network and regular expression method.
[0011] Preferably, in step S1, the pre-processed maritime safety information is encoded by a set encoding rule, and the encoding rule includes:
[0012] (1) Information source: used to describe the maritime agency source of the data, occupying 5 bits;
[0013] (2) Broadcasting station: used to describe the organization that performs the broadcast of maritime distress safety information, occupying 4 bits;
[0014] (3) Information code: used to describe the serial number and year. The serial number consists of 4 digits and is numbered year by year starting from 0001 for different information sources. The year is the last two digits of the current year and occupies 21 bits.
[0015] (4) Alert level: used to describe the urgency of maritime distress safety information, occupying 3 bits;
[0016] (5) Message type: used to describe the category of the maritime distress safety message sent, occupying 4 bits;
[0017] (6) Message subtype: used to describe the subtype of the category to which the maritime distress safety message belongs, occupying 3 bits;
[0018] (7) Validity date: used to describe the validity period of maritime safety information, occupying 21 bits;
[0019] (8) Action area: includes the number of areas and the type of area. The number of areas is used to describe the number of action areas included in this message and occupies 4 bits. The area type is used to describe the area type and its coordinate data and occupies 3 bits. The coordinate data of different types of areas occupy different bits, and each coordinate point occupies 55 bits.
[0020] Preferably, the identifying and removing of redundant information in the maritime safety information message in step S2 comprises the following steps:
[0021] Step S21: The prediction model accumulates and updates the bit pattern of the input maritime safety information in real time. If the current bit number exceeds the set threshold, step S22 is executed;
[0022] Step S22: The prediction model predicts the probability distribution of the next symbol based on the results of feature extraction and semantic analysis, thereby identifying and reducing redundant information.
[0023] Preferably, in step S22, the network parameters of the prediction model are adjusted using a back propagation algorithm, the deviation between the prediction and the true value is measured using a cross entropy loss function, and the weight and bias of each layer are adjusted using an Adam optimizer to optimize the parameters of the prediction model.
[0024] Preferably, step S3 comprises the following steps:
[0025] Step S31: the number of successor characters contained in the current node, the distribution of occurrence of the successor characters, the total frequency of all successor characters of the current node and the frequency information of the successor characters are used as nodes of the context index tree, a root node is created, and the compact maritime safety information message is traversed character by character;
[0026] Step S32: during the traversal process, for the character i in the compact maritime safety information message, determine whether the character i exists in the child node of the current node, if so, execute step S33, otherwise execute step S34;
[0027] Step S33: create a new node for character i, update the frequency information of the subsequent character of the current node, and encode character i according to the frequency information of the current node;
[0028] Step S34: If it does not exist, mark the character i as an escape character, go back to the previous level of the current node, and encode the character i according to the frequency of the escape character.
[0029] Step S35: repeating steps S32 to S34 until all characters in the compact maritime safety information message are processed;
[0030] Step S35: Encode the characters into a bit stream through an adaptive arithmetic encoder.
[0031] Preferably, after returning to the previous level of the current node in step S33, the probability of the escape character is calculated based on the secondary escape estimation SEE, and the escape character is encoded according to the predicted probability distribution of the escape character.
[0032] The present invention also provides a maritime safety information lossless compression system based on context semantic coding, which is implemented based on the above-mentioned maritime safety information lossless compression method based on context semantic coding, and includes: a data cleaning module, a data compact mapping module, a semantic-driven encoding module and a context-aware compression module;
[0033] The data cleaning module is used to clean and pre-process the original maritime safety information, including removing redundant characters, unifying the data format and structuring the information;
[0034] The data compact mapping module converts the cleaned maritime safety information into a compact coding representation through field mapping and structure simplification strategies to ensure that the data is compatible with the format specification of the Beidou short message interface control file;
[0035] The semantic-driven encoding module: performs in-depth semantic analysis and feature extraction on the compact coding representation, captures the potential semantic connection of information by learning the contextual relationship between data, and further compresses the original compact coding into a final coding representation with more concise expression;
[0036] The context-aware compression module constructs a multi-level context index tree, performs dynamic probability prediction on characters based on context, optimizes and compresses escaped characters through adaptive arithmetic coding, and finally generates an efficient bit stream output.
[0037] Preferably, the semantically driven encoding module uses a long short-term memory network (LSTM) to perform deep semantic analysis and feature extraction on the compact coding representation.
[0038] Preferably, the data cleaning module cleans the maritime safety information text, deletes characters irrelevant to the text content, removes irrelevant punctuation marks at the beginning of the maritime safety information text, and performs basic English grammar checks to improve text consistency.
[0039] The benefits of the present invention include at least:
[0040] 1. Use word frequency analysis to build a vocabulary so that the compression algorithm can be optimized for words that appear more frequently in maritime safety information, thereby improving the compression rate;
[0041] 2. By training the prediction model for feature extraction and semantic analysis, it is possible to more accurately identify and remove redundant information, further improve compression efficiency, and generate compact maritime safety information messages that maintain the original information content while reducing the amount of data for easier transmission;
[0042] 3. Context information is used for dynamic probability prediction and adaptive coding to achieve lossless compression, ensuring that the compressed data can be completely restored to its original state. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 A schematic diagram of a method flow of an embodiment of the present invention;
[0044] Figure 2 A structural diagram of the maritime safety information message broadcast content according to an embodiment of the present invention;
[0045] Figure 3 A schematic diagram of a flow chart of semantic-driven coding according to an embodiment of the present invention;
[0046] Figure 4 A schematic diagram of a process of context-aware compression according to an embodiment of the present invention;
[0047] Figure 5 A schematic diagram of a flow chart of an interval encoder used in context-aware compression according to an embodiment of the present invention;
[0048] Figure 6 Schematic diagram of the system framework of an embodiment of the present invention. DETAILED DESCRIPTION
[0049] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the protection scope of the present invention.
[0050] like Figure 1 As shown, an embodiment of the present invention provides a method for lossless compression of maritime safety information based on context semantic coding, comprising the following steps:
[0051] Step S1: Clean and pre-process the original maritime distress safety information, including removing redundant characters, unifying the data format, and structuring the information.
[0052] By standardizing the consistency of information format and content, we ensure that the data meets the transmission requirements of the Beidou short message interface, and provide standardized data input for subsequent encoding and compression operations, thereby improving the accuracy and efficiency of the compression process.
[0053] Specifically, first, text preprocessing is performed on maritime distress safety information, including text cleaning and grammar checking. This process deletes characters irrelevant to the text content, such as spaces, Tab keys, line breaks \n, \r, etc., and removes irrelevant punctuation marks at the beginning of the main body of the maritime distress safety information. At the same time, basic English grammar checking is carried out to improve text consistency.
[0054] After text cleaning and grammar checking are completed, a trained neural network model and regular expressions are used to extract the key content of the maritime distress safety information - the header part of the message. This neural network model is constructed based on historical maritime distress safety information to identify specific key information items, such as fields like information source, dissemination station, information encoding, alert level, etc. For example, for the information "Yunnan Navigational Warning 0005 / 22", the system can parse and obtain that the "information source" is "Yunnan Maritime Safety Administration", the "dissemination station" is "Guangzhou Dissemination Station", and the "information number" is "0005 / 22" and other key information.
[0055] After in-depth analysis of the structural characteristics of maritime distress safety information, it is found that the information source and information number have high regularity in format. Therefore, in natural language processing, for these fields with a fixed structure, the system uses a simple and efficient pattern matching method for information extraction. The regular matching method based on the keyword range first locks the specified range of the target string, and then performs regular expression matching within this range, significantly improving the accuracy of the matching and avoiding redundant matching. For example, for the navigational warning "Zhejiang Navigational Warning 0492, the actual use of weapons training in the East China Sea has ended, cancel the warning of Zhejiang Navigational Warning 0485, please pay attention to all ships.", the system parses according to the following process:
[0056] (1) Design the regular expression R of the keyword A To extract the word "Zhejiang Navigational Warning".
[0057] (2) Use the regular expression R A For matching, take the first matching result in the obtained result set. It can be known that the starting character position is 0, so F = {0}.
[0058] (3) Since the maximum number of characters of the information number is 7, the search range x is appropriately set to 14, and a substring "Zhejiang Navigational Warning 0492, the actual use in the East China Sea" is obtained, so S children = {"Zhejiang Navigational Warning 0492, the actual use in the East China Sea"}.
[0059] (4) Use the regular expression R of the task target B For matching in the above substring to extract the information number.
[0060] (5) Obtain the target matching result "0492".
[0061] For other data items in non-fixed formats, such as alarm level, information type, and affected area, the present invention uses a trained hierarchical multi-label classification network model based on MSML-BERT for information matching and extraction. For example, for the text "The National Marine Forecasting Center issued a blue storm surge alert in accordance with the Marine Disaster Emergency Plan. It is expected that from the afternoon of November 2 to the morning of November 3, there will be a 20 to 50 cm storm surge in the northern and eastern coastal areas of Hainan Island." The module can identify and extract key information such as "Broadcasting Station" is "National Marine Forecasting Center", "Information Type" is "Storm Surge", "Alarm Level" is "Blue Alert", "Valid Date" is "November 2 afternoon to 3 morning", and "Affected Area" is "Northern and Eastern Coast of Hainan Island".
[0062] Step S2: Through field mapping and structure simplification strategies, the standardized information is converted into a compact coding representation to ensure that the data is compatible with the format specification of the Beidou short message interface control file, and to generate a maritime safety information message that complies with the Beidou short message transmission standard.
[0063] Specifically, key information is encoded into a concise representation based on a specific field dictionary, thereby reducing redundant data, and the byte occupancy of information is greatly reduced by optimizing the length and structure of information fields, providing more compressed input for subsequent semantically driven encoding.
[0064] The telegram of Beidou maritime safety information service adopts the short message communication function of Beidou satellite navigation system as the carrier, takes maritime safety information as the data source, and converts the data source into a compact mapping representation format that meets the requirements of Beidou short message data segment through the compact mapping encoding method set by the present invention, so as to ensure efficient transmission.
[0065] like Figure 2 As shown, in the business rules specified in the embodiment of the present invention, the encoding scheme of each data item is as follows:
[0066] The information source occupies 5 bits and refers to the source type of maritime distress safety information. It is used to describe the maritime agency source of the data, such as the Maritime Safety Administration, Ocean Administration or Meteorological Observatory.
[0067] The broadcasting station occupies 4 bits and refers to the agency that carries out the broadcasting of maritime distress safety information, mainly including the National Marine Forecasting Center, Tianjin Broadcasting Station, Shanghai Broadcasting Station and Guangzhou Broadcasting Station.
[0068] The information code occupies 21 bits and the coding format is "serial number + year". The serial number consists of 4 digits and is numbered from 0001 year by year according to the issuing agency; the "year" is the last two digits of the current year.
[0069] The alert level occupies 3 bits and is used to indicate the urgency of maritime distress safety information. For navigation warnings, there are three levels: routine, important and extremely important; for weather warnings, there are four levels: blue, yellow, orange and red; for weather forecasts, there are four levels: slight, general, severe and dangerous.
[0070] The message type occupies 4 bits and is used to indicate the category of the maritime distress safety message being sent.
[0071] The message subtype occupies 3 bits and is used to indicate the subtype of the maritime distress safety message category.
[0072] The validity date occupies 21 bits and is used to indicate the validity period of maritime safety information. When the information content does not specify the cancellation time, the default setting is to cancel it 1 hour after the event subject disappears or is completed; if only the date is indicated in the information, the cancellation time is set to 24:00 on the same day; if there is no clear validity period, all coding bits are set to 0.
[0073] The effective area includes the area quantity code and the area type code. The area quantity code occupies 4 bits and indicates the number of effective areas contained in the message. The area type code occupies 3 bits and is used to indicate the area type and its coordinate data. Different types of area coordinate data codes occupy different bits. The longitude and latitude of a single coordinate point are expressed in the format of "degrees, minutes, seconds" and are accurate to two decimal places after the second. Each coordinate point code occupies 55 bits.
[0074] The maritime safety information data items are of indefinite length. The metadata items of the maritime safety information messages are replaced by dictionary encoding with compact mapping representation to reduce the space occupied by the text and achieve more efficient data transmission.
[0075] Step S3: Use the long short-term memory network LSTM to perform deep semantic analysis and feature extraction on the compact code representation, and further compress the original compact code into a final code representation with more concise expression.
[0076] The LSTM network captures the potential semantic connections of information by learning the contextual relationships between data. It can automatically identify and remove residual redundant information, thereby effectively improving the compression rate. It also reduces the risk of data loss during decoding and information restoration through the sequence prediction capabilities of the LSTM network.
[0077] Specifically, Figure 3As shown in the figure, the maritime safety information is first preprocessed, including steps such as text segmentation and data cleaning, so as to generate standardized input data for subsequent encoding. A vocabulary is built according to the data characteristics to initialize the predictor. The predictor is trained based on the symbol frequency of the input data to generate a coding model, and reads the input data byte by byte to convert it into a bit stream. The semantic-driven encoding module performs deep semantic analysis and feature extraction on the text data through the LSTM network to generate a more compact and expressive representation so that it can be combined with the context model and interval encoder for efficient compression in the future. After the encoding is completed, the data is written to the output file in the form of a bit stream, and the coding rate and related statistical information are recorded.
[0078] During the encoding process, the LSTM network retains the contextual information in the text by virtue of its ability to capture long-term dependencies and understand the semantic relationship between words. Through this semantically driven representation, the module can generate a more compact encoding form than the original text, removing redundant data while retaining key semantic and structural features. The encoded representation processed by the LSTM network is further processed using traditional compression algorithms, which can significantly improve compression efficiency.
[0079] During the byte encoding process, the statistical analyzer in the module performs word frequency analysis on the preprocessed data to determine the distribution characteristics of the words in the data. The vocabulary table built based on the statistical analysis results contains all the input symbols required by the predictor when making predictions, which helps to identify the characteristic patterns of the input data and improve the accuracy of prediction and encoding efficiency.
[0080] The working mechanism of the predictor includes two main steps: "prediction" and "update". In the "prediction" stage, the system updates the input bit sequence by accumulating bit patterns. When the number of bits exceeds the set threshold, the system calls the prediction method. At this time, the system loops through each layer of neurons and performs forward propagation in each layer, copies the output of the current layer to the input buffer of the next layer, calculates the neuron value of the output layer, and indexes and normalizes the output value to make it conform to the probability distribution. In the "update" stage, the system performs backpropagation for each time step, uses the cross entropy loss function to measure the deviation between the prediction and the true value, and uses the Adam optimizer to adjust the weights and biases of each layer to optimize the parameters of the prediction model.
[0081] Step S4: By constructing a multi-level context index tree, each level of the compressed data stream and its frequency information are accurately recorded. During the compression process, the probability of the compressed characters is modeled and predicted based on the context index tree model. For unmatched escape characters, the module efficiently encodes them into bit stream output through dynamic probability estimation and adaptive arithmetic encoder. The following steps are included:
[0082] Step S41: The number of successor characters contained in the current node, the distribution of the occurrence of the successor characters, the total frequency of all successor characters of the current node and the frequency information of the successor characters are used as nodes of the context index tree, a root node is created, and the compact maritime safety information message is traversed character by character.
[0083] Step S42: During the traversal process, for the character i in the compact maritime safety information message, determine whether the character i exists in the child nodes of the current node. If so, execute step S43; otherwise, execute step S44.
[0084] Step S43: Create a new node for the character i, update the frequency information of the subsequent character of the current node, and encode the character i according to the frequency information of the current node.
[0085] Step S44: If it does not exist, mark the character i as an escape character, go back to the previous level of the current node, and encode the character i according to the frequency of the escape character.
[0086] Step S45: Repeat steps S42 to S44 until all characters in the compact maritime safety information message are processed.
[0087] Step S45: Encode the characters into a bit stream through an adaptive arithmetic encoder.
[0088] like Figure 4 As shown, the context-aware compression module matches and escapes the input characters layer by layer through the query module (I) and the query module (II). After the data is input, the query module (I) first performs character matching: if the match is successful at the current level, it is encoded according to the character frequency, and the context tree is updated to continue processing the next character; if the match fails, it is marked as an escape character, and the system encodes based on the frequency information of the escape character, and falls back to the previous level, and continues the matching operation through the query module (II). When the order reaches -1, that is, when an escape occurs at the top level, the compression process ends. The difference between the query module (I) and the query module (II) lies in the different methods for predicting escape characters: the query module (I) uses a fixed escape character frequency, that is, the frequency is 1, while the query module (II) calculates the probability of escape characters based on the quadratic escape estimation (SEE).
[0089] The escape module effectively predicts the probability of escape characters by establishing a simplified escape context model. Compared with the traditional PPM algorithm, this context model and prediction mechanism are more concise and efficient, avoiding the complexity of multi-level prediction. During the prediction process, the escape module only needs to query a single layer of context information to calculate the probability of escape characters based on the total frequency, number of characters, hierarchical information, and parent / child context information of the current context, thereby improving prediction efficiency.
[0090] In addition, the prediction update module passes the prediction probability information to Figure 5 In the encoder module shown in FIG. 1 , the encoder uses an interval coding algorithm to convert probability data into a corresponding bit stream and generates a compression result through an output interface. During the encoding process, the PPMd algorithm generates two types of context probability information: one for binary context, that is, probability information with only one successor character, and the other for multi-context, that is, probability information with more than one successor character. The context-aware compression module adaptively adjusts the coding interval according to the actual context structure, thereby achieving accurate and efficient compression.
[0091] like Figure 6 As shown, an embodiment of the present invention also provides a maritime safety information lossless compression system based on context semantic coding, which is implemented based on the above-mentioned maritime safety information lossless compression method based on context semantic coding, and includes: a data cleaning module, a data compact mapping module, a semantically driven encoding module and a context-aware compression module.
[0092] The data cleaning module is used to clean and pre-process the original maritime safety information, including removing redundant characters, unifying data formats, and structuring the information;
[0093] The data compact mapping module converts the cleaned maritime safety information into a compact coding representation through field mapping and structure simplification strategies, ensuring that the data is compatible with the format specification of the Beidou short message interface control file;
[0094] The semantic-driven encoding module performs in-depth semantic analysis and feature extraction on the compact coding representation. By learning the contextual relationship between data, it captures the potential semantic connection of information and further compresses the original compact coding into a final coding representation with more concise expression.
[0095] The context-aware compression module builds a multi-level context index tree, performs dynamic probability prediction on characters based on the context, optimizes and compresses escaped characters through adaptive arithmetic coding, and finally generates an efficient bitstream output.
[0096] In the present invention, in view of the small text characteristics of maritime safety information, in order to quantitatively evaluate the performance of the compression algorithm, the famous Calgary Corpus was selected as a public data set for testing, and the compression ratio (CR) was used as the main evaluation indicator. Referring to Tables 1 and 2, all evaluation programs are set to extreme compression ratio parameters: lz4-9, gzip--best, xz-9-e, zstd--ultra-22, and brotli-q 11. Although the LZ4 compressor focuses mainly on compression speed rather than compression ratio, it is included in this comparison due to its wide application. The xz compressor is implemented using the LZMA method, while the x3 compressor is implemented using an adaptive context model arithmetic encoder. The comparison results show that the compression algorithm proposed in the present invention, as well as Brotli and xz, perform particularly well in compression performance. In addition, in all test scenarios, the compression algorithm proposed in the present invention is superior to LZ4, gzip, zstd, x3, and PPMd.
[0097] As shown in Table 1, nine different types of text are represented, and many types have more than one representative in order to confirm that the performance of the scheme is consistent for any given type. Ordinary English, both fiction and nonfiction, is represented by two books and two papers, labeled book1, book2, paper1, paper2. More unusual styles of English writing are found in bibliographies (bib) and a collection of unedited news articles (news). Three computer programs represent artificial languages (progc, progl, progp). A recording of a terminal session (trans) is included to show the speed increase that can be achieved by applying compression to slow terminal lines. All of the files mentioned so far use ASCII encoding. A few non-ASCII files are also included: two executable code files (obj1, obj2), some geophysical data (geo), and a bitmap black-and-white picture (pic). The file geo is particularly difficult to compress because it contains a wide range of data values, while the file pic is highly compressible because there are a lot of white space in the picture, represented by long strings of zeros. More details on the individual texts are provided in the books mentioned above. Both the books and the papers give the results of compression experiments on these texts.
[0098] Table 1
[0099]
[0100]
[0101] Table 2 shows the evaluation results on the maritime distress safety information dataset produced by the present invention, using the average compression ratio as the performance evaluation index. This comparative experiment selected Adaptive Arithmetic Coder (AAC), x3, xz, Brotli, ppm and PPMd as benchmark algorithms. AAC is a general lossless compression algorithm, x3 is an efficient dictionary compression algorithm, and xz, Brotli and PPMd have good performance on small public datasets.
[0102] Table 2 Dataset evaluation results
[0103]
[0104] The results in Table 2 show that the algorithm proposed in the present invention achieves significantly better compression rates than other algorithms on the six maritime data subsets, and is only slightly lower than Brotli on the navigation warning dataset, which fully demonstrates the effectiveness of the algorithm in compressing the maritime distress safety information dataset.
[0105] These experimental results not only verify the effectiveness of the algorithm of the present invention, but also show that when performing the compression task of a single small sample of maritime safety information in actual scenarios, the method proposed in the present invention can achieve excellent compression rate and compression efficiency.
[0106] Based on the above description, a specific implementation example of the present invention is given.
[0107] For an original maritime safety message such as "The National Marine Forecasting Center has issued a blue storm surge alert in accordance with the Marine Disaster Emergency Plan. Affected by the cold air, it is expected that during the day on October 16, Bohai Bay and Laizhou Bay will experience a storm surge of 30 to 120 cm. The storm surge warning level in Weifang City, Shandong Province is blue.", the text is first preprocessed to remove the redundant spaces. The standardized text is as follows: "The National Marine Forecasting Center has issued a blue storm surge alert in accordance with the Marine Disaster Emergency Plan. Affected by the cold air, it is expected that during the day on October 16, Bohai Bay and Laizhou Bay will experience a storm surge of 30 to 120 cm. The storm surge warning level in Weifang City, Shandong Province is blue." On this basis, the system uses the trained neural network model and regular expressions to automatically identify and extract key information in the maritime safety information, especially the header part of the message.
[0108] This module can effectively identify and extract the following key information: the source of information is "Oceanic Bureau system organization"; the "broadcast station" is "National Ocean Forecast Center"; the information type is "storm surge"; the alert level is "blue alert"; the effective date is "October 16 daytime"; the effective area is "Bohai Bay and Laizhou Bay". This information is accurately extracted in an automated way, laying the foundation for subsequent processing.
[0109] After obtaining the standardized maritime safety information and extracted key information, the data compact mapping module converts it into a compact mapping representation format that meets the requirements of the Beidou short message data segment. For the specific conversion process, please refer to Table 3, Table 4 and Table 5. Through this module, the original maritime safety information will be converted into the Beidou short message compact mapping format that complies with the Beidou short message transmission protocol. In the example of this article, the converted compact mapping format is:
[0110] “10101000000000000000000001011000110001001000101000000000000000010010”.
[0111] In this process, the text content of maritime safety information will be processed in UTF-8 encoding, and then the semantic-driven encoding module will use the trained model to automatically identify and remove residual redundant information, further improve the compression rate and compress the information into a more concise final encoding representation. The core function of this module is to optimize information expression through deep semantic analysis to achieve more efficient data compression.
[0112] Finally, after being processed by the context-aware compression module, the data stream will pass through a multi-level context index tree structure to record each level and its frequency information. During the compression process, the probability of occurrence of each character to be compressed is modeled and predicted based on the context index tree model. For unmatched escape characters, the module efficiently encodes them through dynamic probability estimation and adaptive arithmetic encoder, and outputs the result as a bit stream. In this way, the module can significantly improve the compression performance and ensure the efficiency and accuracy of the compressed output.
[0113] Finally, by combining the compact mapping coding and text compression parts of the Beidou short message, the final coding is generated in binary data form, which contains all compressed information and is suitable for efficient short message transmission.
[0114] Table 3
[0115]
[0116] Table 4
[0117]
[0118]
[0119] Table 5
[0120] Coded fields Area Type 1 Regional coordinate data Number of bits 3 Unfixed length Original content Sea Area Bohai Bay and Laizhou Bay Encoded content 0 18 Coding results 000 10010
[0121] The present invention proposes a lossless compression method for maritime safety information based on context semantic coding. Through a data cleaning and standardization module and a data compact mapping module, various types of maritime safety information can be encoded and transmitted under the same framework, thereby improving the consistency and stability of transmission; by introducing a deep learning model in the semantically driven coding module, the present invention achieves highly compact expression in the information coding process, eliminates redundant data and greatly improves compression efficiency, and effectively supports the limited data transmission requirements of Beidou short messages; through the collaborative work of an escape prediction algorithm based on context modeling and an adaptive arithmetic encoder, the compression process is optimized, and the context is predicted and dynamically adapted in real time during compression, effectively reducing the number of coding bits, and achieving improved compression rate and decoding efficiency.
[0122] In summary, the method of the present invention deeply combines semantic coding technology with context prediction algorithm, so that maritime distress safety information can achieve significantly improved compression ratio and efficient transmission on the basis of lossless compression, thus meeting the needs of fast and accurate transmission of maritime information.
[0123] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. Only the preferred embodiments of the present invention are expressed. The description is more specific and detailed, but it cannot be understood as limiting the scope of the present invention. As long as there is no contradiction in the combination of these technical features, they should be considered as within the scope of this specification.
[0124] It should be noted that, for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these modifications and improvements all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the attached claims.
Claims
1. A method for lossless compression of maritime safety information based on contextual semantic coding, characterized in that: The following steps are involved: Step S1: pre-processing the original maritime safety information, encoding the pre-processed maritime safety information according to the short message interface control file of the Beidou satellite area and the data format of the maritime safety information, and generating a maritime safety information message; Step S2: performing a word frequency analysis on the pre-processed maritime safety information, constructing a maritime safety information vocabulary according to the word frequency analysis result, using the vocabulary to train a prediction model, using the trained prediction model to perform feature extraction and semantic analysis on the maritime safety information message, identifying and removing redundant information in the maritime safety information message, and obtaining a compact maritime safety information message; Step S3: Using context information, dynamically probabilistically predict and adaptively encode characters in the compact maritime safety information message to achieve lossless compression.
2. The method for lossless compression of maritime safety information based on contextual semantic coding according to claim 1 is characterized in that: The preprocessing of the original maritime safety information in step S1 includes: performing text cleaning and grammar checking on the original maritime safety information, converting the cleaned maritime safety information into a standard format of maritime safety information content, and extracting data of key components of the original maritime information from the format-converted maritime safety information based on a trained neural network and regular expression method.
3. The method for lossless compression of maritime safety information based on contextual semantic coding according to claim 1 is characterized in that: In step S1, the pre-processed maritime safety information is encoded by a set encoding rule, and the encoding rule includes: (1) Information source: used to describe the maritime agency source of the data, occupying 5 bits; (2) Broadcasting station: used to describe the organization that performs the broadcast of maritime distress safety information, occupying 4 bits; (3) Information code: used to describe the serial number and year. The serial number consists of 4 digits and is numbered year by year starting from 0001 for different information sources. The year is the last two digits of the current year and occupies 21 bits. (4) Alert level: used to describe the urgency of maritime distress safety information, occupying 3 bits; (5) Message type: used to describe the category of the maritime distress safety message sent, occupying 4 bits; (6) Message subtype: used to describe the subtype of the category to which the maritime distress safety message belongs, occupying 3 bits; (7) Validity date: used to describe the validity period of maritime safety information, occupying 21 bits; (8) Action area: includes the number of areas and the type of area. The number of areas is used to describe the number of action areas included in this message and occupies 4 bits. The area type is used to describe the area type and its coordinate data and occupies 3 bits. The coordinate data of different types of areas occupy different bits, and each coordinate point occupies 55 bits.
4. The method for lossless compression of maritime safety information based on contextual semantic coding according to claim 1, characterized in that: The step S2 of identifying and removing redundant information in the maritime safety information message comprises the following steps: Step S21: The prediction model accumulates and updates the bit pattern of the input maritime safety information in real time. If the current bit number exceeds the set threshold, step S22 is executed; Step S22: The prediction model predicts the probability distribution of the next symbol based on the results of feature extraction and semantic analysis, thereby identifying and reducing redundant information.
5. The method for lossless compression of maritime safety information based on contextual semantic coding according to claim 4 is characterized in that: In step S22, the network parameters of the prediction model are adjusted using the back propagation algorithm, the deviation between the prediction and the true value is measured using the cross entropy loss function, and the weight and bias of each layer are adjusted using the Adam optimizer to optimize the parameters of the prediction model.
6. The method for lossless compression of maritime safety information based on contextual semantic coding according to claim 1, characterized in that: Step S3 includes the following steps: Step S31: the number of successor characters contained in the current node, the distribution of occurrence of the successor characters, the total frequency of all successor characters of the current node and the frequency information of the successor characters are used as nodes of the context index tree, a root node is created, and the compact maritime safety information message is traversed character by character; Step S32: during the traversal process, for the character i in the compact maritime safety information message, determine whether the character i exists in the child node of the current node, if so, execute step S33, otherwise execute step S34; Step S33: create a new node for character i, update the frequency information of the subsequent character of the current node, and encode character i according to the frequency information of the current node; Step S34: if it does not exist, mark the character i as an escape character, go back to the previous level of the current node, and encode the character i according to the frequency of the escape character; Step S35: repeating steps S32 to S34 until all characters in the compact maritime safety information message are processed; Step S35: Encode the characters into a bit stream through an adaptive arithmetic encoder.
7. The method for lossless compression of maritime safety information based on contextual semantic coding according to claim 6 is characterized in that: After returning to the previous level of the current node in step S33, the probability of the escape character is calculated based on the secondary escape estimation SEE, and the escape character is encoded according to the predicted probability distribution of the escape character.
8. A maritime safety information lossless compression system based on context semantic coding, implemented based on a maritime safety information lossless compression method based on context semantic coding as claimed in any one of claims 1 to 7, characterized in that: include: Data cleaning module, data compact mapping module, semantic-driven encoding module, and context-aware compression module; The data cleaning module is used to clean and pre-process the original maritime safety information, including removing redundant characters, unifying the data format and structuring the information; The data compact mapping module converts the cleaned maritime safety information into a compact coding representation through field mapping and structure simplification strategies to ensure that the data is compatible with the format specification of the Beidou short message interface control file; The semantic-driven encoding module: performs in-depth semantic analysis and feature extraction on the compact coding representation, captures the potential semantic connection of information by learning the contextual relationship between data, and further compresses the original compact coding into a final coding representation with more concise expression; The context-aware compression module constructs a multi-level context index tree, performs dynamic probability prediction on characters based on context, optimizes and compresses escaped characters through adaptive arithmetic coding, and finally generates an efficient bit stream output.
9. A maritime safety information lossless compression system based on contextual semantic coding according to claim 8, characterized in that: The semantically driven encoding module uses a long short-term memory network (LSTM) to perform in-depth semantic analysis and feature extraction on compact coding representation.
10. A maritime safety information lossless compression system based on contextual semantic coding according to claim 8, characterized in that: The data cleaning module cleans the maritime safety information text, deletes characters irrelevant to the text content, removes irrelevant punctuation marks at the beginning of the maritime safety information text, and performs basic English grammar checks to improve text consistency.
Citation Information
Patent Citations
Remote channel message compression method and system for electricity consumption collection system
CN105553625A
Maritime safety information efficient lossless compression system and method based on Beidou short message
CN117459069A