Data Compression Method Based on Shipbuilding Mobile Management Platform
Through characteristic analysis and adaptive data segmentation, the appropriate compression algorithm is selected to classify and compress ship construction data, and embedded it into the mobile management platform, solving the problems of time-consuming and waste of resources in ship construction data transmission, and achieving efficient and convenient data management and use.
Patent Information
- Application Number
- CN202510360756.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-03-26
AI Technical Summary
Uncompressed ship construction data takes a long time during transmission, network bandwidth limitations lead to inefficient transmission, increased computer processing burden, and waste resources and inefficient traditional data processing methods.
Through feature analysis, adaptive data segmentation and classification processing, appropriate compression algorithms are selected to classify and compress ship construction data, and embedded in the mobile management platform for real-time collection, compression, transmission and storage. Staff can view and process the compressed data at any time.
Without losing data integrity, it significantly reduces storage and transmission costs, improves data processing efficiency and accuracy, ensures data security and reliability, and improves user experience and management convenience.
Smart Images

Figure CN119883138B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a data compression method based on a shipbuilding mobile management platform. Background Art
[0002] Shipbuilding is a complex project involving multiple fields, covering the entire process from ship design to actual construction. In this process, designers determine the structure, size, and performance parameters of the ship based on factors such as the ship's purpose, navigation conditions, and customer requirements. Subsequently, the construction team will use materials such as steel and aluminum alloy according to these design parameters, and gradually complete the construction of various parts such as the hull, ship's shell, and ship's frame through processes such as welding, cutting, and forming.
[0003] During the shipbuilding process, a large amount of data is generated. This data covers information on various aspects from raw material procurement, production progress, quality control to final delivery. This data is crucial for shipbuilding enterprises. It not only helps monitor the production progress, optimize the production process, but also provides strong support for quality control and after-sales service.
[0004] However, the following problems exist in actual operations:
[0005] Uncompressed shipbuilding data may be very time-consuming during transmission. Especially when the data volume is huge, the limitation of network bandwidth may lead to low transmission efficiency. A large amount of uncompressed data will increase the burden on the computer to process data. Traditional data processing methods may not be able to fully utilize system resources, resulting in resource waste or low processing efficiency. Summary of the Invention
[0006] The purpose of the present invention is to provide a data compression method based on a shipbuilding mobile management platform to solve the problems raised in the above background art.
[0007] To achieve the above purpose, the present invention provides the following technical solution: A data compression method based on a shipbuilding mobile management platform, including the following steps:
[0008] Characteristic analysis, obtaining shipbuilding data, performing data analysis on the shipbuilding data, preprocessing the shipbuilding data, and caching the shipbuilding data according to the data type after preprocessing;
[0009] Compression algorithm selection, based on the analysis results of the characteristic analysis, selecting different compression algorithms for different types of shipbuilding data;
[0010] Adaptive data segmentation, which segments shipbuilding data. According to the data type and the amount of data, it adaptively segments the shipbuilding data into several subsets. Among them, when processing text data, it is segmented based on paragraphs, and when processing image and video data, it is segmented based on key frames;
[0011] Data classification processing, which classifies data according to the characteristic analysis, compression algorithm selection, and the data type and compression requirements obtained after adaptive data segmentation, and generates different compression strategies based on the data categories respectively;
[0012] Data compression, which constructs a data compression module and performs compression processing on the data subsets in the shipbuilding data based on the compression strategy, and stores the compressed data in the compressed data storage module. When performing data compression, it uses multi-threaded processing to simultaneously process the compression and decompression tasks of multiple data subsets;
[0013] Embedding into the mobile management platform, which embeds the data compression module into the shipbuilding mobile management platform. Through the mobile management platform, it realizes the real-time collection, compression, transmission, storage, and decompression of shipbuilding data. Staff can view and process the compressed data at any time through the mobile management platform.
[0014] Further, the characteristic analysis specifically further includes the following steps:
[0015] Collect shipbuilding data from all aspects of shipbuilding, and distinguish the shipbuilding data into key data and non-key data. Among them, the key data includes design drawings and process parameters, and the non-key data includes other data except design drawings and process parameters;
[0016] Conduct type analysis on the collected non-key data, and based on the data type, visually display the proportion, distribution, and storage characteristics of the non-key data through a data visualization tool. Among them, the data types include text data, image data, video data, and structured data, and the visual display is carried out by drawing one or more of bar charts, pie charts, and graphs. At the same time, conduct word frequency statistics and keyword extraction on the text data through a text analysis tool;
[0017] Analyze the relevance between various aspects of the shipbuilding data, identify redundant information in the shipbuilding data, and at the same time, use data deduplication and compression algorithms to identify and remove the redundant information;
[0018] Perform classified caching on key data and non-key data based on the data type of the shipbuilding data.
[0019] Further, the compression algorithm selection specifically further includes the following steps:
[0020] Select different compression algorithms based on the type of shipbuilding data;
[0021] Among them, for critical data, a lossless compression algorithm is adopted, and the lossless compression algorithm is one of the LZ77 compression algorithm and the Huffman coding compression algorithm;
[0022] For non-critical data, a lossy compression algorithm is adopted. The lossy compression algorithm includes Huffman coding compression algorithm, H.264 compression algorithm, and columnar storage. Among them, the Huffman coding compression algorithm is selected when processing text data; the H.264 compression algorithm is selected when processing image data and video data; columnar storage is selected when processing structured data.
[0023] Furthermore, the adaptive data segmentation specifically further includes the following steps:
[0024] Identify paragraphs in the text data. Among them, when identifying paragraphs, it is achieved by identifying line breaks, indents, and paragraph marks in the text;
[0025] Parse the content of each paragraph in the text data. Among them, when parsing the content, it is achieved by keyword extraction and theme recognition;
[0026] Based on the length of the paragraphs in the text data and the size of the entire text data volume, determine the segmentation points, and based on the segmentation points, segment the text data in the shipbuilding data into several subsets;
[0027] Decode and convert the format of the video data, extract the key frames in the video, and perform content evaluation on the extracted key frames. Among them, when performing content evaluation, it is achieved by image recognition and object detection;
[0028] According to the content and data volume size of the key frames, segment the video data in the shipbuilding data into several subsets. Among them, each subset contains a series of key frames and some video frames between the key frames.
[0029] Furthermore, the adaptive data segmentation specifically further includes the following steps:
[0030] For large shipbuilding data files and data streams, perform block compression on the subsets after data segmentation;
[0031] For small shipbuilding data files and short data segments, perform overall compression. At the same time, dynamically adjust the segmentation strategy of each subset of the shipbuilding data blocks according to the real-time nature and importance of the data.
[0032] Furthermore, the data classification processing specifically further includes the following steps:
[0033] Determine an appropriate compression level based on the importance of the data and the storage space requirements. Among them, for data with high importance and the need to save storage space, select a high compression level. Conversely, a low compression level can be selected;
[0034] Set corresponding compression parameters according to the requirements of the selected compression algorithm. Among them, the compression parameters include block size and dictionary size;
[0035] Create corresponding compression strategies based on the classification of the data and the compression parameters.
[0036] Furthermore, the data compression specifically further includes the following steps:
[0037] Take each data subset as an independent compression task and allocate it to different threads for processing;
[0038] Maintain synchronization and communication between threads in multi-threaded processing, and dynamically adjust the number of threads according to system resources and the amount of compression tasks;
[0039] After compression is completed, store the compressed data in a compressed data storage module, where the compressed data storage module is one of a disk and a database;
[0040] When storing the compressed data, assign a unique identifier to each compressed data subset.
[0041] Furthermore, the embedding of the mobile management platform specifically further includes the following steps:
[0042] Embed the data compression module into the mobile management platform through the API interface method;
[0043] Set a real-time data acquisition function in the mobile management platform to collect data during the shipbuilding process in real time. The collected shipbuilding data is automatically transmitted to the data compression module for compression processing after characteristic analysis, compression algorithm selection, adaptive data segmentation, and data classification processing;
[0044] The compressed shipbuilding data is transmitted and stored through the mobile management platform, and the compressed shipbuilding data is encrypted during the transmission process;
[0045] Staff can view and process the compressed data at any time through the mobile management platform. When data needs to be viewed, the platform will automatically decompress the compressed data and display it on the user interface.
[0046] Furthermore, analyze the correlation between each link in the shipbuilding data, including:
[0047] Select a pair of linked process data from the shipbuilding data as the first process data and the second process data;
[0048] Extract the characteristic indicators of the first process data and calculate the first reliability coefficient of the characteristic indicators in the set of characteristic attributes:
[0049]
[0050] Wherein, K 1 is the first reliability coefficient; b a is the characteristic indicator; F 1 is the membership function of the characteristic indicator in the set of characteristic attributes A ; U 1 is the domain where the characteristic indicator is located; n is the number of all indicators in the set of characteristic attributes A ; b x is the x-th indicator in the set of characteristic attributes A; T (b a , b x ) is the matching degree between the characteristic indicator and the x-th indicator in the set of characteristic attributes A;
[0051] Determine the matching indicators related to the characteristic indicators in the second process data according to the association relationship between the first process data and the second process data;
[0052] Calculate the second reliability coefficient of the matching indicators in the relative set of characteristic attributes:
[0053]
[0054] Wherein, K 2 is the second reliability coefficient; b c is the matching indicator; F 2 is the membership function of the matching indicator in the relative set of characteristic attributes (R-A) ; m is the number of all indicators in the relative set of characteristic attributes (R-A) ; b y is the y-th indicator in the relative set of characteristic attributes (R-A) ; T (b c , b y ) is the matching degree between the matching indicator and the y-th indicator in the relative set of characteristic attributes (R-A) ; U 2 is the domain where the matching indicator is located; R is the total set of characteristic attributes;
[0055] Based on the first reliability coefficient and the second reliability coefficient, the third reliability coefficient of the correlation relationship between the first link data and the second link data:
[0056]
[0057] Wherein, K 3 is the third reliability coefficient;
[0058] Compare the third reliability coefficient with a preset threshold, and regard the correlation relationship with the third reliability coefficient greater than or equal to the preset threshold as a valid relationship, and eliminate the correlation relationship with the third reliability coefficient less than the preset threshold.
[0059] Furthermore, preprocess the shipbuilding data, including:
[0060] Extract strings from the shipbuilding data to determine the key strings;
[0061] Establish a pattern matching automaton according to the pattern set and the regular pattern string;
[0062] Input the key strings into the pattern matching automaton for initial matching with the regular pattern string, and divide the regular pattern string into multiple regular substrings based on preset rules; the multiple regular substrings are respectively matched with the key strings;
[0063] Determine the special substrings in the key strings that do not match the regular substrings, match the special substrings with the abnormal strings in the abnormal database, and determine the abnormal data in the shipbuilding data according to the matching information;
[0064] Determine the partial data in the shipbuilding data that is associated with the abnormal data;
[0065] Perform smoothing processing and data repair processing on the partial data based on the least squares method to determine the overall data;
[0066] Determine the mapping data corresponding to the abnormal data in the overall data according to the correlation relationship between the abnormal data and the partial data;
[0067] Perform interpolation calculation on the mapping data and perform sequential identification to determine the first identification data sequence;
[0068] Perform interpolation calculation on the abnormal data and perform sequential identification to determine the second identification data sequence;
[0069] Determine at least 3 matching points between the first identification data sequence and the second identification data sequence, including the first matching point and the last matching point;
[0070] Establish the correspondence between the first identification data sequence and the second identification data sequence according to the start matching point and the end matching point, and repair the abnormal data according to the correspondence to obtain the repaired shipbuilding data.
[0071] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0072] 1. Through the characteristic analysis of shipbuilding data, the present invention selects a compression algorithm according to the data characteristics, creates an adaptive data segmentation strategy according to the data characteristics, classifies different types of data for further improving the compression efficiency. At the same time, it designs an efficient data compression and decompression module and embeds it into the mobile management platform. Staff can view and process the compressed data through the mobile management platform at any time to ensure the real-time compression and decompression of data, thereby significantly reducing the storage and transmission costs without losing the data integrity.
[0073] 2. By analyzing the characteristics of shipbuilding data, the present invention can more accurately understand the essence and characteristics of the data, which helps to select appropriate compression algorithms and strategies according to the data characteristics in the subsequent data processing process, thereby improving the efficiency and accuracy of data processing. Classifying the data according to the data type and compression requirements and generating different compression strategies can further ensure the security and reliability of the data. For the data that needs to be frequently accessed and modified, an algorithm that is fast but has a slightly lower compression ratio is used to ensure the real-time performance and availability of the data; for the data that does not need to be frequently accessed but needs to be stored for a long time, an algorithm with a higher compression ratio but a slightly slower decompression speed is used to ensure the long-term storage and security of the data.
[0074] 3. The present invention performs block compression on the segmented subsets, which can effectively reduce the storage space occupied by the data. For large shipbuilding data files and data streams, this block compression method can further reduce the storage cost while maintaining the integrity and accessibility of the data. By dynamically adjusting the segmentation strategy of each subset of the data block, different processing can be performed according to the real-time performance and importance of the data. For real-time data that requires quick response, smaller segmentation units and a lower compression ratio can be used to ensure the timeliness and accuracy of the data; for data with lower importance or non-real-time data, larger segmentation units and a higher compression ratio can be used to further save storage space.
[0075] 4. In the present invention, the staff can view and process the compressed data at any time through the mobile management platform without performing additional data decompression or conversion operations. This not only improves the user experience but also reduces the operation difficulty and complexity, making the management and use of data more convenient. By performing real-time data acquisition, compression, transmission, storage, and decompression on the data through the mobile management platform, the continuity and traceability of the data can be ensured, enabling the enterprise to better track the source and change process of the data and providing strong support for subsequent data analysis and decision-making. Brief Description of the Drawings
[0076] Figure 1 It is a schematic flow chart of data compression for the mobile management platform of shipbuilding in the present invention. Detailed Embodiments
[0077] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0078] To solve the technical problems that uncompressed shipbuilding data may be very time-consuming during transmission, especially when the data volume is huge, the network bandwidth limitation may lead to low transmission efficiency, a large amount of uncompressed data will increase the burden on the computer to process data, and traditional data processing methods may not fully utilize system resources, resulting in resource waste or low processing efficiency, please refer to Figure 1 , the present invention provides the following technical solutions:
[0079] A data compression method based on the mobile management platform of shipbuilding includes the following steps:
[0080] Characteristic analysis: Obtain shipbuilding data, perform data analysis on the shipbuilding data, preprocess the shipbuilding data, and cache the shipbuilding data according to the data type after preprocessing;
[0081] Compression algorithm selection: Based on the analysis results of the characteristic analysis, select different compression algorithms for different types of shipbuilding data;
[0082] Adaptive data segmentation: Segment the shipbuilding data, and adaptively segment the shipbuilding data into several subsets according to the data type and data volume. Among them, when processing text data, it is segmented based on paragraphs, and when processing image and video data, it is segmented based on key frames;
[0083] Data classification and processing: According to the data types and compression requirements obtained through characteristic analysis, compression algorithm selection, and adaptive data segmentation, the data is divided into different categories, and different compression strategies are generated based on the data categories respectively.
[0084] Data compression: Build a data compression module and perform compression processing on the data subsets in the shipbuilding data based on the compression strategy, and store the compressed data in the compressed data storage module. When performing data compression, multi-threading is used to handle the compression and decompression tasks of multiple data subsets simultaneously.
[0085] Embed into the mobile management platform: Embed the data compression module into the shipbuilding mobile management platform, and through the mobile management platform, perform real-time collection, compression, transmission, storage, and decompression of the shipbuilding data. Staff can view and process the compressed data through the mobile management platform at any time.
[0086] Specifically, through the characteristic analysis of the shipbuilding data, select the compression algorithm according to the data characteristics, create an adaptive data segmentation strategy according to the data characteristics, classify different types of data for further improving the compression efficiency. At the same time, design an efficient data compression and decompression module and embed it into the mobile management platform to ensure the real-time compression and decompression of data, thereby significantly reducing the storage and transmission costs without losing data integrity.
[0087] In the above embodiment, through characteristic analysis and compression algorithm selection, the most suitable compression algorithm can be selected for different types of shipbuilding data, thereby improving the data processing efficiency. At the same time, the adaptive data segmentation technology can divide the data into several subsets according to the data type and data volume, further improving the flexibility and efficiency of data processing.
[0088] In the above embodiment, embedding the data compression module into the shipbuilding mobile management platform enables staff to view and process the compressed data through the mobile management platform at any time, greatly improving the practicality and convenience of the mobile management platform, and making the data management in the shipbuilding process more efficient and flexible.
[0089] Characteristic analysis: Obtain the shipbuilding data, perform data analysis on the shipbuilding data, preprocess the shipbuilding data, and cache the shipbuilding data according to the data type after preprocessing.
[0090] Characteristic analysis specifically further includes the following steps:
[0091] Collect shipbuilding data from all aspects of shipbuilding, and distinguish the shipbuilding data into critical data and non-critical data. Among them, the critical data includes design drawings and process parameters, and the non-critical data includes other data except design drawings and process parameters;
[0092] Conduct a type analysis on the collected non-critical data, and based on the data type, use a data visualization tool to visually display the proportion, distribution, and storage characteristics of the non-critical data. Among them, the data type includes text data, image data, video data, and structured data, and the visual display is carried out by drawing one or more of bar charts, pie charts, and graphs. At the same time, use a text analysis tool to conduct word frequency statistics and keyword extraction on the text data;
[0093] Analyze the correlation between each link in the shipbuilding data, identify redundant information in the shipbuilding data, and at the same time, use data deduplication and compression algorithms to identify and remove the redundant information;
[0094] Perform classified caching of critical data and non-critical data based on the data type of the shipbuilding data.
[0095] In the above embodiment, by analyzing the characteristics of the shipbuilding data, including distinguishing critical data from non-critical data, analyzing data types, data proportions, distributions, and storage characteristics, etc., the essence and characteristics of the data can be understood more accurately, which helps to select appropriate compression algorithms and strategies according to the characteristics of the data in the subsequent data processing process, thereby improving the efficiency and accuracy of data processing.
[0096] In the above embodiment, by visually displaying and conducting word frequency statistics and other analyses on the non-critical data, the distribution of the data can be understood more intuitively, and then redundant information can be identified and removed. This can not only reduce the occupation of storage space and storage costs, but also improve the efficiency of data transmission, reduce transmission time and network bandwidth consumption.
[0097] In the above embodiment, classifying and processing the data according to the data type and compression requirements, and generating different compression strategies can further ensure the security and reliability of the data. For data that needs to be frequently accessed and modified, an algorithm that is fast but has a slightly lower compression ratio can be used to ensure the real-time performance and availability of the data; for data that does not need to be frequently accessed but needs to be stored for a long time, an algorithm with a higher compression ratio but a slightly slower decompression speed can be used to ensure the long-term storage and security of the data.
[0098] Compression algorithm selection: Based on the analysis results of the characteristic analysis, select different compression algorithms for different types of shipbuilding data;
[0099] Compression algorithm selection, which specifically includes the following steps:
[0100] Select different compression algorithms based on the types of shipbuilding data;
[0101] Among them, for critical data, a lossless compression algorithm is adopted, and the lossless compression algorithm is one of the LZ77 compression algorithm and the Huffman coding compression algorithm;
[0102] For non-critical data, a lossy compression algorithm is adopted. The lossy compression algorithm includes Huffman coding compression algorithm, H.264 compression algorithm, and columnar storage. Among them, the Huffman coding compression algorithm is selected when processing text data; the H.264 compression algorithm is selected when processing image data and video data; columnar storage is selected when processing structured data.
[0103] In the above embodiments, for critical data such as design drawings and process parameters, a lossless compression algorithm such as the LZ77 compression algorithm or the Huffman coding compression algorithm is selected, which can ensure that the integrity of the data is not lost during the compression process, and ensure that the critical data still maintains its original accuracy and accuracy after compression, providing reliable data support for subsequent shipbuilding work. For non-critical data, a lossy compression algorithm such as the Huffman coding compression algorithm, the H.264 compression algorithm, and columnar storage can reduce the storage space occupied by the data to a certain extent, and achieve compression by removing redundant information in the data or reducing the accuracy of the data, thereby reducing the storage cost.
[0104] In the above embodiments, selecting a suitable compression algorithm for different types of shipbuilding data can make the data more efficient during the processing process. Through reasonable data compression, the burden of data transmission and processing can be reduced, thereby improving the response speed and performance of the mobile management platform, which is crucial for real-time monitoring and remote management of the shipbuilding process, and helps to improve management efficiency and decision-making level.
[0105] Adaptive data segmentation, segment the shipbuilding data, and adaptively segment the shipbuilding data into several subsets according to the data type and the size of the data volume. Among them, when processing text data, it is segmented based on paragraphs, and when processing image and video data, it is segmented based on key frames;
[0106] Adaptive data segmentation, which specifically includes the following steps:
[0107] Identify the paragraphs in the text data. Among them, when identifying paragraphs, it is achieved by identifying line breaks, indents, and paragraph marks in the text;
[0108] Parse the content of each paragraph in the text data. Among them, when parsing the content, it is achieved by keyword extraction and theme recognition;
[0109] Based on the length of paragraphs in the text data and the size of the entire text data, determine the splitting points, and split the text data in the shipbuilding data into several subsets based on the splitting points;
[0110] Decode and convert the format of the video data, extract the key frames in the video, and perform content evaluation on the extracted key frames, where the content evaluation is achieved through image recognition and object detection;
[0111] According to the content and size of the key frames, split the video data in the shipbuilding data into several subsets, where each subset contains a series of key frames and some video frames between the key frames.
[0112] For large shipbuilding data files and data streams, perform block compression on the subsets after data splitting;
[0113] For small shipbuilding data files and short data segments, perform overall compression. At the same time, dynamically adjust the splitting strategy of each subset of the shipbuilding data blocks according to the real-time nature and importance of the data.
[0114] In the above embodiment, by identifying the paragraphs in the text data and the key frames in the video data, this solution can adaptively split the shipbuilding data into several subsets according to the data type and size, enabling subsequent data processing and analysis to be carried out on smaller and more concentrated data units, thereby improving the processing efficiency. Performing block compression on the split subsets can effectively reduce the storage space occupied by the data. For large shipbuilding data files and data streams, this block compression method can further reduce the storage cost while maintaining the integrity and accessibility of the data.
[0115] In the above embodiment, by dynamically adjusting the splitting strategy of each subset of the data blocks, different processing can be performed according to the real-time nature and importance of the data. For real-time data that requires quick response, smaller splitting units and lower compression ratios can be used to ensure the timeliness and accuracy of the data; for data with lower importance or non-real-time data, larger splitting units and higher compression ratios can be used to further save storage space.
[0116] Data classification processing, according to the data types and compression requirements obtained from characteristic analysis, compression algorithm selection, and adaptive data splitting, classify the data into different categories, and generate different compression strategies based on the data categories;
[0117] Data classification processing specifically further includes the following steps:
[0118] Determine an appropriate compression level according to the importance of the data and the storage space requirements. Among them, for data with high importance and the need to save storage space, select a high compression level. Conversely, a low compression level can be selected.
[0119] Set corresponding compression parameters according to the requirements of the selected compression algorithm. Among them, the compression parameters include block size and dictionary size.
[0120] Create corresponding compression strategies based on the classification of the data and the compression parameters.
[0121] In the above embodiments, by classifying the data, different compression strategies can be formulated according to the characteristics and compression requirements of different types of data, ensuring that important data can be compressed more efficiently, while avoiding unnecessary compression losses, thus meeting the diverse data processing requirements in the shipbuilding process. According to the importance of the data and the storage space requirements, determine an appropriate compression level, which can minimize the storage space occupancy. For data with high importance and the need to save storage space, adopting a high compression level can significantly reduce the data volume and improve the storage and transmission efficiency; while for data with relatively low importance or that does not require excessive compression, selecting a low compression level can ensure the integrity and readability of the data.
[0122] In the above embodiments, according to the requirements of the selected compression algorithm, set corresponding compression parameters such as block size and dictionary size, which can further optimize the compression quality and efficiency. By adjusting these parameters, the characteristics of different types of data can be adapted, improving the adaptability and flexibility of the compression algorithm. By creating corresponding compression strategies based on the classification of the data and the compression parameters, the data management can be made more targeted. Different data categories adopt different compression strategies, which helps to better manage and utilize data resources in the shipbuilding process, improving the management efficiency and decision-making level.
[0123] For data compression, construct a data compression module and perform compression processing on data subsets in shipbuilding data based on the compression strategy, and store the compressed data in the compressed data storage module. When performing data compression, use multi-threaded processing to simultaneously handle the compression and decompression tasks of multiple data subsets.
[0124] Data compression specifically further includes the following steps:
[0125] Take each data subset as an independent compression task and allocate it to different threads for processing.
[0126] Maintain synchronization and communication between threads in multi-threaded processing, and dynamically adjust the number of threads according to system resources and the amount of compression tasks.
[0127] After compression, the compressed data is stored in a compressed data storage module, where the compressed data storage module is either a disk or a database;
[0128] When storing the compressed data, a unique identifier is assigned to each subset of the compressed data.
[0129] In the above embodiment, by treating each data subset as an independent compression task and distributing it to different threads for processing, parallel compression can be achieved, making full use of computing resources, greatly accelerating the compression processing speed, improving the overall compression efficiency, making the compression of large-scale shipbuilding data more efficient, maintaining synchronization and communication between threads in multi-threaded processing, and dynamically adjusting the number of threads according to system resources and the amount of compression tasks, which can ensure the reasonable allocation and efficient utilization of system resources. This avoids waste of resources and also ensures the smooth progress of the compression task.
[0130] In the above embodiment, after compression, the compressed data is stored in a compressed data storage module, which can be either a disk or a database, providing flexibility for data storage and subsequent use. At the same time, a unique identifier is assigned to each subset of the compressed data, facilitating data identification, retrieval, and management, and improving the convenience of data management.
[0131] In the above embodiment, through multi-threaded processing and dynamic resource adjustment, the stability and reliability of the data compression process can be ensured. Even when facing a large number of data compression tasks, the system can still operate efficiently and stably, reducing the risk of system crashes or performance degradation caused by compression tasks. Due to the improvement of compression efficiency and the optimized utilization of resources, the mobile management platform can respond more quickly when processing user requests, reducing the waiting time of users, helping to improve the user experience, and enhancing user satisfaction and trust in the platform.
[0132] Embed the mobile management platform, embed the data compression module into the shipbuilding mobile management platform, and through the real-time collection, compression, transmission, storage, and decompression of shipbuilding data by the mobile management platform, staff can view and process the compressed data at any time through the mobile management platform;
[0133] Embedding the mobile management platform specifically further includes the following steps:
[0134] Embed the data compression module into the mobile management platform through the API interface method;
[0135] Set a real-time data collection function in the mobile management platform to collect data during the shipbuilding process in real time. The collected shipbuilding data is automatically transmitted to the data compression module for compression processing after characteristic analysis, compression algorithm selection, adaptive data segmentation, and data classification processing;
[0136] The compressed shipbuilding data is transmitted and stored through a mobile management platform, and the compressed shipbuilding data is encrypted during the transmission process;
[0137] Staff can view and process the compressed data through the mobile management platform at any time. When viewing the data is required, the platform will automatically decompress the compressed data and display it on the user interface.
[0138] Furthermore, analyze the correlation between each link in the shipbuilding data, including:
[0139] Select a pair of linked link data in the shipbuilding data as the first link data and the second link data;
[0140] Extract the characteristic index of the first link data and calculate the first reliability coefficient of the characteristic index in the characteristic attribute set:
[0141]
[0142] Among them, K 1 is the first reliability coefficient; b a is the characteristic index; F 1 is the membership function of the characteristic index in the characteristic attribute set A ; U 1 is the domain where the characteristic index is located; n is the number of all indexes in the characteristic attribute set A ; b x is the x-th index in the characteristic attribute set A; T (b a , b x ) is the matching degree between the characteristic index and the x-th index in the characteristic attribute set A;
[0143] According to the correlation between the first link data and the second link data, determine the matching index related to the characteristic index in the second link data;
[0144] Calculate the second reliability coefficient of the matching index in the relative characteristic attribute set:
[0145]
[0146] Among them, K 2 is the second reliability coefficient; b c is the matching index; F 2 is the membership function of the matching index in the relative characteristic attribute set (R-A)Membership function on; m is the set of relative characteristic attributes (R-A) The number of all indicators above; b y Is the set of relative characteristic attributes (R-A) The y-th indicator in; T (b c , b y ) is the matching degree between the matching indicator and the set of relative characteristic attributes (R-A) The matching degree of the y-th indicator in; U 2 Is the universe of discourse where the matching indicator is located; R Is the total set of characteristic attributes;
[0147] According to the first reliability coefficient and the second reliability coefficient, the third reliability coefficient of the correlation relationship between the first link data and the second link data:
[0148]
[0149] Among them, K 3 Is the third reliability coefficient;
[0150] Compare the third reliability coefficient with the preset threshold, and regard the correlation relationship with the third reliability coefficient greater than or equal to the preset threshold as the valid relationship, and eliminate the correlation relationship with the third reliability coefficient less than the preset threshold.
[0151] The working principle of the above technical solution: In this embodiment, the shipbuilding data includes data of each link. Select a pair of link data with a correlation relationship in the shipbuilding data as the first link data and the second link data.
[0152] In this embodiment, statistical methods (such as descriptive statistics, data visualization, etc.) are used to explore the distribution, trend and relationship of the first link data. Through EDA, it is possible to initially determine which features may be important and their relationships. Feature selection is performed based on domain knowledge, statistical tests (such as chi-square test, t-test, ANOVA, etc.) or machine learning algorithms (such as random forest, gradient boosting machine, etc.) to determine the most important feature indicators. The set of characteristic attributes describes the information contained in each record in the table, as well as the data type and constraint conditions of this information, and also describes the characteristics or states of each category.
[0153] In this embodiment, the membership function on the set of characteristic attributes refers to a function used to describe the membership degree of a characteristic indicator on the set of characteristic attributes. In fuzzy logic, the membership function is usually used to characterize the degree or probability of an element belonging to different fuzzy sets. In feature engineering, the membership function can be used to evaluate and describe the mutual relationship and attributes between features. The membership function can be a Gaussian membership function:
[0154]
[0155] Among them, x represents the eigenvalue, μ represents the mean value, σ represents the standard deviation; e is the natural constant. By adjusting the values of the mean and the standard deviation, the membership degree of the feature on the feature attribute set can be described. In addition to the Gaussian membership function, there are other different forms of membership functions, such as the triangular membership function, the trapezoidal membership function, etc.
[0156] The universe of discourse is an important concept in mathematical logic, which refers to the set of all individuals or objects involved in a reasoning or discussion. It usually consists of three parts: a non-empty set of elements M', including the basic elements of the universe of discourse. A non-empty set of functions on M', where each function has the Cartesian product of one M' or multiple M's as the domain and M' as the range. A non-empty set of propositions about M', and each proposition represents the logical relationship between the elements of M', between the functions, and between the elements and the functions.
[0157] In this embodiment, the second reliability coefficient of the matching index in the relative feature attribute set is calculated. The relative feature attribute set is the relative set obtained by removing the feature attribute set from the total feature attribute set. The universe of discourse where the matching index is located is the same as the universe of discourse where the feature index is located.
[0158] Beneficial effects of the above technical solution: accurately calculate the first reliability coefficient of the feature index in the feature attribute set and the second reliability coefficient of the matching index in the relative feature attribute set, and then accurately calculate the third reliability coefficient of the correlation relationship between the data of the first link and the data of the second link, improve the accuracy of judging the magnitude relationship between the third reliability coefficient and the preset threshold, and judge whether the correlation relationship between the data of each link is a valid relationship or an invalid relationship, so as to realize the analysis of the correlation between each link in the shipbuilding data and realize the collation of the correlation relationship.
[0159] Further, preprocess the shipbuilding data, including:
[0160] Extract strings from the shipbuilding data to determine the key strings;
[0161] Establish a pattern matching automaton according to the pattern set and the regular pattern string;
[0162] Input the key strings into the pattern matching automaton to perform an initial match with the regular pattern string, and divide the regular pattern string into multiple regular substrings based on a preset rule; the multiple regular substrings are respectively matched with the key strings;
[0163] Determine the special substrings in the key string that do not match the regular substrings, match the special substrings with the abnormal strings in the abnormal database, and determine the abnormal data in the shipbuilding data according to the matching information;
[0164] Determine the partial data in the shipbuilding data that is associated with the abnormal data;
[0165] Perform smoothing processing and data repair processing on the partial data based on the least squares method to determine the overall data;
[0166] Determine the mapping data corresponding to the abnormal data in the overall data according to the association relationship between the abnormal data and the partial data;
[0167] Perform interpolation calculation on the mapping data and perform sequential identification to determine the first identification data sequence;
[0168] Perform interpolation calculation on the abnormal data and perform sequential identification to determine the second identification data sequence;
[0169] Determine at least 3 matching points between the first identification data sequence and the second identification data sequence, including the first matching point and the last matching point;
[0170] Establish the corresponding relationship between the first identification data sequence and the second identification data sequence according to the first matching point and the last matching point, and perform repair processing on the abnormal data according to the corresponding relationship to obtain the repaired shipbuilding data.
[0171] The working principle of the above technical solution: In this embodiment, based on the string extraction technology, key information is extracted from the shipbuilding data.
[0172] In this embodiment, a deterministic finite automaton (DFA) or a non-deterministic finite automaton (NFA) is used to implement pattern matching. Determining the automaton type: First, it is necessary to determine whether to use a DFA or an NFA to construct the pattern matching automaton. A DFA usually has a higher space overhead but faster matching speed; while an NFA has a lower space overhead but may have a slightly slower matching speed. Constructing the state transition diagram: According to the given regular pattern string, construct the state transition diagram of the automaton. Each state represents matching a certain prefix pattern string, and state transitions are made according to different input characters. Determining the accepting states: Mark which states in the automaton are accepting states, that is, the complete pattern string has been matched. State transition function: Define the transition function between states, including the rules for transferring to the next state according to the input character. Matching algorithm: Implement the matching algorithm, that is, perform state transitions character by character according to the input text, and finally determine whether the pattern string has been matched. The preset rule is to divide the minimum substring of each string.
[0173] In this embodiment, after verifying the data that matches the normal substring, no other processing is performed. However, for the special substrings in the key string that do not match the regular substring, it is verified and determined whether they are abnormal data. The special substrings are matched with the abnormal strings in the abnormal database, and based on the matching information, the abnormal data in the shipbuilding data is determined.
[0174] In this embodiment, some data is smoothed and data repaired based on the least squares method to determine the overall data, including: Selecting an appropriate least squares model: In data smoothing and repair processing, the least squares method is used to fit the data and find the best fitting curve. Fitting the data: Using the least squares method to fit some data to obtain a fitting curve or function. The core idea of the least squares method is to adjust the model parameters by minimizing the sum of the squares of the residuals, so that the fitting curve optimally conforms to the data distribution. Data repair processing: Based on the fitting curve or function, data repair processing is performed on missing values or outliers. The missing data or abnormal data is filled or repaired according to the predicted values of the fitting curve, making the overall data more complete and accurate. Determining the overall data: According to the partially processed data and the fitting curve after repair processing, the overall data can be determined. According to the actual situation, the overall data can be inferred and filled through the fitting curve to obtain a complete data set.
[0175] In this embodiment, the mapped data is the correct data part corresponding to the abnormal data, and the amount of data of the mapped data is generally larger than that of the abnormal data.
[0176] In this embodiment, interpolation calculation is performed on the mapped data and sequential identification is carried out to determine the first identified data sequence. Interpolation calculation is performed on the mapped data through an interpolation algorithm to fill in missing data or perform data smoothing processing. Commonly used interpolation algorithms include linear interpolation, polynomial interpolation, spline interpolation, etc. Sequential identification: Sequential identification is performed on the interpolated data, which can be identified in chronological order or other specified orders. Ensure that each data point in the data is correctly identified and arranged in a certain order.
[0177] In this embodiment, the method for performing interpolation calculation on the abnormal data is the same as that for performing interpolation calculation on the mapped data, and will not be elaborated here.
[0178] In this embodiment, based on the first matching point and the last matching point, the corresponding relationship between the first identified data sequence and the second identified data sequence is established, that is, their corresponding relationship is determined. Based on the corresponding relationship, the abnormal data is repaired to obtain the repaired shipbuilding data.
[0179] Advantages of the above technical solution: Based on the pattern matching automaton, special substrings in the key string that do not match the regular substrings are determined, and the special substrings are matched with the abnormal strings in the abnormal database to determine the abnormal data in the shipbuilding data. For the abnormal data, data supplementation processing is performed through the partial data associated with the abnormal data to obtain the overall data. Based on the association relationship between the abnormal data and the partial data, mapping data corresponding to the abnormal data is determined in the overall data; by establishing the corresponding relationship between the first identification data sequence and the second identification data sequence, the abnormal data is repaired according to the corresponding relationship to obtain the repaired shipbuilding data, improving the accuracy of the shipbuilding data.
[0180] In the above embodiment, the staff can view and process the compressed data at any time through the mobile management platform without performing additional data decompression or conversion operations, which not only improves the user experience but also reduces the operation difficulty and complexity, making the management and use of data more convenient. By performing real-time data acquisition, compression, transmission, storage, and decompression through the mobile management platform, the continuity and traceability of the data can be ensured, enabling the enterprise to better track the source and change process of the data and providing strong support for subsequent data analysis and decision-making.
[0181] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered by the protection scope of the present invention.
Claims
1. A data compression method based on a shipbuilding mobile management platform, characterized in that Including the following steps: Characteristic analysis, obtaining shipbuilding data, performing data analysis on the shipbuilding data, preprocessing the shipbuilding data, and caching the shipbuilding data according to the data type after preprocessing; Compression algorithm selection, based on the analysis results of characteristic analysis, selecting different compression algorithms for different types of shipbuilding data; Adaptive data segmentation, segmenting the shipbuilding data, and adaptively segmenting the shipbuilding data into several subsets according to the data type and the size of the data volume, wherein, when processing text data, it is segmented based on paragraphs, and when processing image and video data, it is segmented based on key frames; Data classification processing, according to the data type and compression requirements obtained after characteristic analysis, compression algorithm selection, and adaptive data segmentation, classifying the data into different categories, and generating different compression strategies based on the data categories respectively; Data compression, constructing a data compression module and performing compression processing on the data subsets in the shipbuilding data based on the compression strategy, and storing the compressed data in the compressed data storage module. When performing data compression, multi-threaded processing is used to simultaneously process the compression and decompression tasks of multiple data subsets; Embedding into the mobile management platform, embedding the data compression module into the shipbuilding mobile management platform, and through the mobile management platform, real-time collection, compression, transmission, storage, and decompression of the shipbuilding data. Staff can view and process the compressed data at any time through the mobile management platform; Preprocessing the shipbuilding data, including: Performing string extraction on the shipbuilding data to determine the key strings; Establishing a pattern matching automaton according to the pattern set and the regular pattern string; Inputting the key strings into the pattern matching automaton to perform an initial match with the regular pattern string, and dividing the regular pattern string into multiple regular sub-strings based on preset rules; the multiple regular sub-strings are respectively matched with the key strings; Determining the special sub-strings in the key strings that do not match the regular sub-strings, matching the special sub-strings with the abnormal strings in the abnormal database, and determining the abnormal data in the shipbuilding data according to the matching information; Determining the partial data in the shipbuilding data that is associated with the abnormal data; Performing smoothing processing and data repair processing on the partial data based on the least squares method to determine the overall data; Determining the mapped data corresponding to the abnormal data in the overall data according to the association relationship between the abnormal data and the partial data; Performing interpolation calculation on the mapped data and performing sequential identification to determine the first identification data sequence; Performing interpolation calculation on the abnormal data and performing sequential identification to determine the second identification data sequence; Determining at least 3 matching points between the first identification data sequence and the second identification data sequence, including the first matching point and the last matching point; Establishing the correspondence relationship between the first identification data sequence and the second identification data sequence according to the first matching point and the last matching point, and performing repair processing on the abnormal data according to the correspondence relationship to obtain the repaired shipbuilding data.
2. The data compression method based on the shipbuilding mobile management platform according to claim 1, wherein: The said characteristic analysis specifically further includes the following steps: Collect shipbuilding data from all aspects of shipbuilding, and distinguish the shipbuilding data into critical data and non-critical data. Among them, the critical data includes design drawings and process parameters, and the non-critical data includes other data except design drawings and process parameters; Conduct a type analysis on the collected non-critical data, and based on the data type, use a data visualization tool to visually display the proportion, distribution, and storage characteristics of the non-critical data. Among them, the data types include text data, image data, video data, and structured data, and the visual display is performed by drawing one or more of bar charts, pie charts, and graphs. At the same time, use a text analysis tool to perform word frequency statistics and keyword extraction on the text data; Analyze the correlation between each link in the shipbuilding data, identify redundant information in the shipbuilding data, and at the same time, use data deduplication and compression algorithms to identify and remove redundant information; Perform classified caching on critical data and non-critical data based on the data type of the shipbuilding data.
3. The data compression method based on the shipbuilding mobile management platform according to claim 1, characterized in that: The selection of the compression algorithm specifically further includes the following steps: Select different compression algorithms based on the type of shipbuilding data; Among them, for critical data, a lossless compression algorithm is used, and the lossless compression algorithm is one of the LZ77 compression algorithm and the Huffman coding compression algorithm; For non-critical data, a lossy compression algorithm is used. The lossy compression algorithm includes Huffman coding compression algorithm, H.264 compression algorithm, and columnar storage. Among them, the Huffman coding compression algorithm is selected when processing text data; the H.264 compression algorithm is selected when processing image data and video data; columnar storage is selected when processing structured data.
4. The data compression method based on the shipbuilding mobile management platform according to claim 1, characterized in that: The adaptive data segmentation specifically further includes the following steps: Identify paragraphs in the text data, and when identifying paragraphs, it is achieved by identifying line breaks, indents, and paragraph marks in the text; Parse the content of each paragraph in the text data, and when parsing the content, it is achieved by keyword extraction and theme recognition; Based on the length of the paragraphs in the text data and the size of the entire text data volume, determine the segmentation point, and based on the segmentation point, divide the text data in the shipbuilding data into several subsets; Decode and convert the format of the video data, extract the key frames in the video, and perform content evaluation on the extracted key frames. When performing content evaluation, it is achieved by image recognition and object detection; According to the content and data volume size of the key frames, divide the video data in the shipbuilding data into several subsets, where each subset contains a series of key frames and some video frames between the key frames.
5. The data compression method based on the shipbuilding mobile management platform according to claim 4, characterized in that: The adaptive data segmentation specifically further includes the following steps: For large shipbuilding data files and data streams, perform block compression on the subsets after data segmentation; For small shipbuilding data files and short data segments, perform overall compression. At the same time, dynamically adjust the segmentation strategy of each subset of the shipbuilding data blocks according to the real-time nature and importance of the data.
6. The data compression method based on the shipbuilding mobile management platform according to claim 1, characterized in that: The data classification processing specifically further includes the following steps: Determine an appropriate compression level according to the importance of the data and the demand for storage space. Among them, for data with high importance and the need to save storage space, select a high compression level; otherwise, select a low compression level. Set corresponding compression parameters according to the requirements of the selected compression algorithm. Among them, the compression parameters include block size and dictionary size. Create corresponding compression strategies based on the classification of the data and the compression parameters.
7. The data compression method based on the shipbuilding mobile management platform according to claim 1, characterized in that: The data compression specifically further includes the following steps: Take each data subset as an independent compression task and allocate it to different threads for processing. Maintain synchronization and communication between threads in multi-threaded processing and dynamically adjust the number of threads according to system resources and the amount of compression tasks. After compression is completed, store the compressed data in a compressed data storage module, where the compressed data storage module is one of a disk and a database. When storing the compressed data, assign a unique identifier to each compressed data subset.
8. The data compression method based on the shipbuilding mobile management platform according to claim 1, characterized in that: The embedding into the mobile management platform specifically further includes the following steps: Embed the data compression module into the mobile management platform through the API interface method. Set a real-time data acquisition function in the mobile management platform to collect data during the shipbuilding process in real time. The collected shipbuilding data, after characteristic analysis, compression algorithm selection, adaptive data segmentation, and data classification processing, is automatically transmitted to the data compression module for compression processing. The compressed shipbuilding data is transmitted and stored through the mobile management platform, and the compressed shipbuilding data is encrypted during the transmission process. Staff can view and process the compressed data at any time through the mobile management platform. When data needs to be viewed, the platform will automatically decompress the compressed data and display it on the user interface.
Citation Information
Patent Citations
Data processing method and device
CN113595557A
Data security storage system
CN117857177A