Multi-modal data intelligent filing management system
Through the multimodal data intelligent archiving and management system, the semantic gap problem of multimodal data in cross-platform migration in traditional technologies is solved, the reliable storage and efficient use of data are achieved, and the cross-platform migration integrity and retrieval response speed of data are improved.
Patent Information
- Application Number
- CN202510949440.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-07-10
AI Technical Summary
Traditional multimodal data archiving technology has difficulty in achieving dynamic business-related restoration of data, which affects the integrity of the business context after it leaves the native system and makes it difficult to fully utilize it.
A multimodal data intelligent archiving management system is adopted to achieve trusted storage and cross-platform migration and utilization of data throughout its entire life cycle through data mapping module, attribute tree construction module, differential processing module, data packet generation module and verification module.
It improves the cross-platform migration integrity and retrieval response speed of data, reduces the risk of sensitive information leakage, optimizes storage resource utilization, and improves data utilization efficiency and user interaction experience.
Smart Images

Figure CN120723722A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal data processing, and in particular to a multimodal data intelligent archiving and management system. Background Art
[0002] Traditional archiving technologies primarily utilize a "format conversion + metadata annotation" processing model, which, to a certain extent, enables the archiving and storage of multimodal data. However, this approach lacks the ability to restore dynamic business connections within the data. For example, in government approval systems, the semantic associations between process data (e.g., processing nodes, approval opinions) and attached materials (e.g., scanned documents, on-site videos) require specific business logic analysis. While existing technologies can convert databases to common formats like XML and annotate basic information with metadata, they struggle to fully preserve the complete semantic chain of implicit logical mappings in complex business scenarios (e.g., temporal associations and semantic dependencies between approval nodes and attached materials).
[0003] The document mentions that "it is difficult to extract and restore the semantics of structured discrete data." In actual applications, the historical data of a provident fund business system uses traditional archiving technology and lacks deep mapping of the dynamic semantics of business processes, resulting in difficulty in tracing some business processes. This technical bottleneck makes the integrity of the business context of multimodal data easily affected after it is separated from the native system, thereby restricting the in-depth utilization of the data. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a multimodal data intelligent archiving and management system to achieve trusted evidence storage, cross-platform migration and utilization, and intelligent management and control throughout the data life cycle.
[0005] In order to solve the above technical problems, the technical solutions of the present invention are as follows: In a first aspect, a multimodal data intelligent archiving and management system comprises: The data mapping module is used to extract multimodal data from the native business environment, generate a data unit set, construct a virtual reference plane, establish a reference rectangular grid topology structure and divide the dynamic coordinate axis in the plane, map the data unit set to grid coordinates, and output the grid coordinate mapped data unit set; The attribute tree construction module is used to map the density distribution of the data unit set according to the grid coordinates, construct a multidimensional attribute tree, generate dynamic evaluation nodes, calculate the weight parameter β value of each node, and output the attribute tree structure with β value; A differential processing module is used to perform differential processing on the grid coordinate mapping data unit set according to the β value of each node and the preset classification threshold to obtain a processed data unit set; The data packet generation module is used to convert the processed data unit set into a file set, add a differentiated blockchain evidence identification based on the attribute tree node β value, and output a pre-archived data packet with the β identification; The verification module is used to push data packets to the pre-archiving verification terminal, perform hierarchical verification, output the data packets that have passed the verification, parse the passed data packets and import them into the database, allocate storage resource priorities, and output archived data with storage level identification; The dynamic adjustment module is used to provide retrieval services for archived data, collect user access behavior data in real time, feed the access behavior data back to the attribute tree construction module, and dynamically adjust the coordinate axis division rules.
[0006] Furthermore, multimodal data is extracted from the native business environment to generate a data unit set, and a virtual reference plane is constructed. A reference rectangular grid topology structure is established within the plane and dynamic coordinate axes are divided. The data unit set is mapped to grid coordinates, and a grid coordinate mapping data unit set is output, including: Extract multimodal data from the native business environment and generate independent data units with data type labels, confidentiality level labels, and call frequency labels; Perform preprocessing operations on independent data units, including screening valid data units based on preset rules, performing cleaning operations on invalid fields, sorting sensitive fields according to confidentiality requirements and data units according to call frequency, and outputting a set of preprocessed data units; Taking the preprocessed data unit set as input data, a three-dimensional virtual reference plane is constructed based on the data type label, the confidentiality level label and the frequency label; A benchmark rectangular grid topology structure covering the entire data range is established in the virtual reference plane, and the dynamic coordinate axis scale is divided according to the distribution density of the input data in the three-dimensional feature space. Each preprocessed data unit is mapped to the corresponding coordinate unit of the grid, and the grid coordinate mapping data unit set is output.
[0007] Furthermore, a reference rectangular grid topology structure covering the entire data range is established in the virtual reference plane, and the dynamic coordinate axis scale is divided according to the distribution density of the input data in the three-dimensional feature space. Each preprocessed data unit is mapped to the corresponding coordinate unit of the grid, and a grid coordinate mapping data unit set is output, including: Based on the three-dimensional space coordinate range of the pre-processed data unit set in terms of data type, confidentiality level, and call frequency, a rectangular grid structure that completely covers all data coordinates is generated; According to the distribution density of each preprocessed data unit in the three-dimensional space, the scale interval of each coordinate axis is adaptively adjusted, and each preprocessed data unit is positioned to the corresponding coordinate unit in the grid according to the data type label value, confidentiality level label value, and call frequency label value; Integrate all data units that have completed coordinate positioning to form a grid coordinate mapping data unit set.
[0008] Furthermore, based on the density distribution of the grid coordinate mapping data unit set, a multidimensional attribute tree is constructed, dynamic evaluation nodes are generated, and the weight parameter β value of each node is calculated, and the attribute tree structure with β value is output, including: Map the data unit set according to the grid coordinates, and count the number of data units in each grid unit as the density value; A multidimensional attribute tree is constructed based on the number of data units in each grid cell as the density value, where each tree node corresponds to a grid cell and the node level is determined by the position of the grid cell in the three-dimensional feature space; Generate a dynamic evaluation node for each tree node, obtain the density value of the grid unit corresponding to the corresponding node, and calculate the characteristic association strength value of the data unit in the corresponding grid unit based on the association relationship between the three types of labels: data type, confidentiality level and call frequency; According to the density value of the corresponding grid unit of the corresponding node and the characteristic correlation strength of the data unit in the corresponding grid unit, the node weight parameter β value is obtained, and a multidimensional attribute tree structure carrying the β value of each node is output.
[0009] Furthermore, according to the β value of each node and the preset classification threshold, the grid coordinate mapping data unit set is subjected to differential processing to obtain a processed data unit set, including: Read the β value of each node in the attribute tree structure with β value, and divide the node level according to the preset threshold range, including: defining nodes whose β value is greater than or equal to a first preset threshold as high β value nodes; A node whose β value is greater than or equal to the second preset threshold and less than the first preset threshold is defined as a medium β value node; defining a node whose β value is less than a second preset threshold as a low β value node; Differentiated processing is performed on the grid coordinate mapping data unit set, including enhancing the desensitization strength of sensitive fields for data units associated with high β value nodes; increasing the frequency of data cleaning execution for data units associated with medium β value nodes; and reducing the data sorting priority weight for data units associated with low β value nodes, and outputting the optimized data unit set.
[0010] Furthermore, the processed data unit set is converted into a file set, and a differentiated blockchain evidence identification is added based on the attribute tree node β value, and a pre-archived data package with the β identification is output, including: Converting the optimized data unit set into a structured file set; Generate differentiated blockchain identifiers based on the attribute tree node β value, specifically including: Strong verification and evidence identification is added to the files corresponding to high β value nodes; Add standard evidence identification to the corresponding files of the medium β value nodes; Add basic evidence identification to the corresponding files of low β value nodes; Attach a corresponding identifier to each structured file and output a pre-archived data package with a β identifier.
[0011] Furthermore, the data packets are pushed to the pre-archiving verification terminal, where hierarchical verification is performed, and the verified data packets are output. The passed data packets are parsed and imported into the database, storage resource priorities are allocated, and archived data with storage level identification is output, including: The pre-archived data package with the β mark is pushed to the pre-archived verification terminal. The verification process of the corresponding strength is triggered according to the β mark level. If the verification passes, the parsing and import are performed to obtain the parsed data; The parsed data is assigned storage resource priority based on the β identification value, imported into the storage database, and a storage level identification corresponding to the β value is attached to the archived data entering the database, and the archived data with the storage level identification is output.
[0012] In a second aspect, a computing device includes: one or more processors; The storage device is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the system.
[0013] According to a third aspect, a computer-readable storage medium stores a program, which implements the system when executed by a processor.
[0014] The above solution of the present invention includes at least the following beneficial effects: Converting multimodal data into a standardized grid coordinate system resolves the semantic gap between native data and archival systems, enabling independent parsing of data outside of its native environment and improving integrity during cross-platform migration. Dynamic coordinate axes adaptively adjust based on data density, improving coordinate accuracy and query response speed in high-frequency / high-sensitivity data areas. Beta values are calculated based on density distribution and label correlation strength to quantify the business value and sensitivity of data. A multidimensional attribute tree structure explicitly illustrates data distribution logic, allowing business personnel to intuitively understand the sources of data importance (e.g., high density or high confidentiality level). High-beta data undergoes enhanced desensitization (e.g., full-field encryption), significantly reducing the risk of sensitive information leakage and meeting high-level security requirements. Medium-beta data undergoes increased scrubbing frequency, reducing integrity error rates. Low-beta data is prioritized, freeing up computing resources and improving query speed for high-frequency data. Differentiated blockchain evidence identification is added based on beta values. High-value data is dual-signed for immutability, while low-value data is simplified for on-chain storage costs. Standardized file set encapsulation addresses heterogeneous multimodal data formats and improves cross-system archiving compatibility.
[0015] High-beta data is fully validated; processes for medium and low-beta data are streamlined, improving overall validation efficiency. High-beta data is stored on SSDs, reducing access latency, while low-beta data is archived for storage, effectively reducing costs and achieving "high-speed access to hot data and low-cost storage for cold data." Real-time collection of user access behavior automatically optimizes coordinate axis partitioning rules, continuously improving the retrieval hit rate of frequently queried data and enhancing the user interaction experience. An adaptive learning mechanism allows data classification logic to evolve with business needs, reducing the cost of manual rule maintenance. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a schematic diagram of a multimodal data intelligent archiving and management system provided by an embodiment of the present invention.
[0017] Figure 2 This is a flow chart of an embodiment of the present invention, which provides a method for mapping the density distribution of data units according to grid coordinates, constructing a multidimensional attribute tree, generating dynamic evaluation nodes, calculating the weight parameter β value of each node, and outputting an attribute tree structure with β value. DETAILED DESCRIPTION
[0018] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0019] like Figure 1As shown, an embodiment of the present invention provides a multimodal data intelligent archiving and management system, including: The data mapping module is used to extract multimodal data from the native business environment, generate a data unit set, construct a virtual reference plane, establish a reference rectangular grid topology structure and divide the dynamic coordinate axis in the plane, map the data unit set to grid coordinates, and output the grid coordinate mapped data unit set; The attribute tree construction module is used to map the density distribution of the data unit set according to the grid coordinates, construct a multidimensional attribute tree, generate dynamic evaluation nodes, calculate the weight parameter β value of each node, and output the attribute tree structure with β value; A differential processing module is used to perform differential processing on the grid coordinate mapping data unit set according to the β value of each node and the preset classification threshold to obtain a processed data unit set; The data packet generation module is used to convert the processed data unit set into a file set, add a differentiated blockchain evidence identification based on the attribute tree node β value, and output a pre-archived data packet with the β identification; The verification module is used to push data packets to the pre-archiving verification terminal, perform hierarchical verification, output the data packets that have passed the verification, parse the passed data packets and import them into the database, allocate storage resource priorities, and output archived data with storage level identification; The dynamic adjustment module is used to provide retrieval services for archived data, collect user access behavior data in real time, feed the access behavior data back to the attribute tree construction module, and dynamically adjust the coordinate axis division rules.
[0020] In an embodiment of the present invention, a structured mapping of multimodal data from a native environment to standardized grid coordinates is achieved, thereby improving the independent survival capability of data outside the native system and cross-platform semantic consistency. A multidimensional attribute tree is constructed through data density distribution and the weight β value is calculated to provide a quantitative basis for data importance assessment and intelligent classification. Based on the β value, hierarchical desensitization, cleaning and sorting optimization are performed on the data to improve the utilization efficiency of high-frequency data while ensuring the security of sensitive data. Differentiated evidence identification is added in combination with blockchain technology to achieve tamper-proof evidence and operation traceability throughout the data life cycle. Through a hierarchical verification mechanism and dynamic allocation of storage resources, data quality is ensured and storage efficiency is optimized, and high-frequency and high-value data are preferentially allocated to high-speed storage media.
[0021] In a preferred embodiment of the present invention, extracting multimodal data from a native business environment, generating a data unit set, constructing a virtual reference plane, establishing a reference rectangular grid topology structure and dividing dynamic coordinate axes within the plane, mapping the data unit set to grid coordinates, and outputting a grid coordinate mapping data unit set may include: Extract multimodal data from the native business environment and generate independent data units with data type labels, confidentiality level labels, and call frequency labels; Perform preprocessing operations on independent data units, including screening valid data units based on preset rules, performing cleaning operations on invalid fields, sorting sensitive fields according to confidentiality requirements and data units according to call frequency, and outputting a set of preprocessed data units; Taking the preprocessed data unit set as input data, a three-dimensional virtual reference plane is constructed based on the data type label, the confidentiality level label and the frequency label; A reference rectangular grid topology covering the entire data range is established in the virtual reference plane. The dynamic coordinate axis scale is divided according to the distribution density of the input data in the three-dimensional feature space. Each preprocessed data unit is mapped to the corresponding coordinate unit of the grid, and a grid coordinate mapping data unit set is output, which specifically includes: Based on the three-dimensional space coordinate range of the pre-processed data unit set in terms of data type, confidentiality level, and call frequency, a rectangular grid structure that completely covers all data coordinates is generated; According to the distribution density of each preprocessed data unit in the three-dimensional space, the scale interval of each coordinate axis is adaptively adjusted, and each preprocessed data unit is positioned to the corresponding coordinate unit in the grid according to the data type label value, confidentiality level label value, and call frequency label value; Integrate all data units that have completed coordinate positioning to form a grid coordinate mapping data unit set.
[0022] In an embodiment of the present invention, multimodal data, such as text, images, audio, and structured database tables, is collected from native business systems (such as government approval systems and financial transaction platforms) using APIs or data interface adapters to achieve full data capture. For example, application form text, on-site inspection photos, approval voice recordings, and process database tables from government approval processes are extracted. Categories (such as "text," "image," and "structured data") are defined based on data format. For example, ID card scans are labeled "image" and approval opinion text is labeled "text." Classification levels (such as public, internal, secret, and confidential) are assigned based on business rules. For example, medical records involving personal privacy are labeled "secret," and public policy documents are labeled "public." Data is statistically accessed over the past 12 months and categorized as high-frequency (greater than 100 times), medium-frequency (10-100 times), and low-frequency (less than 10 times). For example, frequently accessed user basic information tables are labeled "high-frequency," while less frequently accessed historical archives are labeled "low-frequency."
[0023] Independent data unit preprocessing: Filter invalid data based on preset rules (such as data integrity thresholds and business relevance indicators). For example, eliminate records with missing key fields (such as blank ID numbers) or exclude units that have exceeded the business validity period (such as temporary approval data from three years ago). Identify and process outliers (such as garbled characters "###" in text fields) and duplicate values (such as duplicate records of the same approval process), and standardize them through string matching or clustering algorithms. For example, replace non-numeric characters in the "ID number" field with spaces. Perform differentiated desensitization according to confidentiality level: Public data: not desensitized (such as policy document titles); Internal data: Partially desensitized (e.g., the middle 8 digits of the ID card number are replaced with "****"); Confidential data: Full field encryption (e.g. medical diagnosis records encrypted using the AES algorithm).
[0024] Data unit sorting: Sort by call frequency in descending order, prioritizing high-frequency data. For example, frequently accessed user transaction records are placed at the front of the preprocessing queue, while low-frequency historical logs are placed at the back.
[0025] 3D virtual reference plane construction: X-axis: data type (e.g., 0-100, where 0 represents text, 100 represents images, and intermediate values represent mixed types); Y-axis: confidentiality level (0-100, 0 is public, 100 is confidential, linearly mapped according to the confidentiality level weight); Z-axis: Call frequency (0-100, 0 is low frequency, 100 is high frequency, converted to coordinate values using a logarithmic function to avoid high-frequency data being concentrated at the top).
[0026] Normalize the label value of each preprocessed data unit to the range [0, 100]. For example, an image file (type label value 80), confidentiality level (confidentiality label value 70), and high-frequency call (frequency label value 90) has a three-dimensional coordinate of (80, 70, 90).
[0027] Baseline rectangular grid topology and dynamic coordinate axis division: Calculate the maximum value (MaxX, MaxY, MaxZ) and minimum value (MinX, MinY, MinZ) of all data cells in three-dimensional space, and generate a minimum rectangular grid containing all coordinates. The range is [MinX-Margin, MaxX+Margin]×[MinY-Margin, MaxY+Margin]×[MinZ-Margin, MaxZ+Margin], where Margin is the expansion threshold (such as 5% of the coordinate range) to ensure that edge data is covered.
[0028] Divide the three-dimensional space into an initial uniform grid (e.g., 10×10×10). Count the number of data cells within each subgrid, defining density as the sum of subgrid data and the total data volume. For high-density areas (density greater than 1.5 times the mean), reduce the scale interval (e.g., from 10 to 5), while for low-density areas (density less than 0.5 times the mean), increase the scale interval (e.g., from 10 to 20). For example, high-frequency confidential image data is concentrated in the (70-90, 60-80, 80-100) areas. Adjust the scale interval on the X, Y, and Z axes in these areas from 10 to 5 to improve coordinate accuracy.
[0029] Data cell positioning and grid mapping: Based on the adjusted dynamic scale, the normalized coordinates (X, Y, Z) of the data cell are converted to grid coordinates (GridX, GridY, GridZ). For example, for a data cell with coordinates (85, 75, 95) in an X-axis scale interval of 5, GridX = (85 - MinX) ÷ 5, rounded to the nearest integer. GridX, GridY, and GridZ are used to locate the specific grid cell. For example, the three-dimensional coordinates (GridX = 17, GridY = 15, GridZ = 19) correspond to the 17th row, 15th column, and 19th level of the grid. If multiple data cells are mapped to the same grid, they are sorted by hash value or timestamp to ensure uniqueness. All grid cells are traversed to check for missing or incorrectly mapped coordinates. The mapping results are stored in a JSON or XML file using the "grid coordinate - data cell ID - label - raw data pointer" structure, forming a searchable mapping dataset. For example, each grid cell stores the index of all data cells in the area, enabling quick query of corresponding data by coordinate.
[0030] Through labeling and grid mapping, heterogeneous data is uniformly converted into a structured coordinate system, resolving the semantic gap between native data and the archive management system and enabling data to be independently parsed outside of the native system. Coordinate axis scales are automatically adjusted based on data density, improving coordinate accuracy in high-frequency / high-sensitivity data areas and reducing storage redundancy. Combining confidentiality level labeling with grid positioning enables regionalized, isolated storage of sensitive data (e.g., confidential data is concentrated in specific grid areas and encrypted separately), reducing the risk of data leakage. Three-dimensional coordinate mapping preserves the relationship between data type, call frequency, and confidentiality level. For example, the grid positioning of high-frequency confidential image data directly reflects its business importance, providing a basis for subsequent data utilization (such as priority archiving and permission control). The standardized grid coordinate system supports lossless data migration between different business systems and hardware environments. For example, when migrating government data from an old approval system to a new archive management system, grid mapping can quickly re-establish data associations.
[0031] In a preferred embodiment of the present invention, a multidimensional attribute tree is constructed based on the density distribution of the grid coordinate mapping data unit set, dynamic evaluation nodes are generated, and the weight parameter β value of each node is calculated. The attribute tree structure with the β value is output, which may include: Map the data unit set according to the grid coordinates, and count the number of data units in each grid unit as the density value; A multidimensional attribute tree is constructed based on the number of data units in each grid cell as the density value, where each tree node corresponds to a grid cell and the node level is determined by the position of the grid cell in the three-dimensional feature space; Generate a dynamic evaluation node for each tree node, obtain the density value of the grid unit corresponding to the corresponding node, and calculate the characteristic association strength value of the data unit in the corresponding grid unit based on the association relationship between the three types of labels: data type, confidentiality level and call frequency; According to the density value of the corresponding grid unit of the corresponding node and the characteristic correlation strength of the data unit in the corresponding grid unit, the node weight parameter β value is obtained, and a multidimensional attribute tree structure carrying the β value of each node is output.
[0032] In an embodiment of the present invention, based on the generated grid coordinate mapping data unit set (such as a 100×100×100 three-dimensional grid), each grid unit (GridX, GridY, GridZ) is traversed in the order of the X, Y, and Z axis coordinates. For example, starting from (0, 0, 0), all grids are accessed row by row, column by column, and layer by layer. For each grid unit, the list of data unit IDs stored in the area is queried by index, and the number of data units is counted as the density value. For example, if the grid (15, 20, 5) contains 200 data units, the density value is recorded as 200. If the grid unit is empty (density value = 0), it is marked as an "empty node" and is treated as a leaf node when the tree structure is subsequently constructed; if the density value exceeds a threshold (such as twice the mean), it is marked as a "high-density area" and the child nodes are constructed first.
[0033] Hierarchical construction of multidimensional attribute tree: The root node represents the entire three-dimensional feature space, covering the coordinate range of all grid cells (e.g., X∈[0, 100], Y∈[0, 100], Z∈[0, 100]). The first-level child nodes divide the root node into several subspaces based on the X-axis coordinate (e.g., X<33, 33≤X<66, X≥66). Each subspace corresponds to a child node, and the node name records the X-axis range (e.g., "X0-33"). The second-level child nodes further divide the first-level child nodes based on the Y-axis coordinate (e.g., Y<33, 33≤Y<66, Y≥66), forming a two-dimensional "X range - Y range" subspace (e.g., "X0-33_Y0-33"). The third-level child nodes divide the root node based on the Z-axis coordinate (e.g., Z<33, 33≤Z<66, Z≥66), ultimately forming three-dimensional subspace nodes (e.g., "X0-33_Y0-33_Z0-33"), each of which corresponds to a specific grid cell.
[0034] Each tree node stores the grid coordinate range it covers (e.g., X∈[a, b], Y∈[c, d], Z∈[e, f]) and is associated with a list of density values for all grid cells within that range. For example, the third-level node "X0-33_Y0-33_Z0-33" corresponds to the first 33% of the 3D grid, encompassing the grid cells from (0, 0, 0) to (33, 33, 33).
[0035] Dynamic evaluation node generation and feature association strength calculation: For each tree node, the density values of all grid cells it covers are aggregated, and a weighted average (e.g., weighted by grid area) is calculated as the node's base density metric. For example, if a node covers five grid cells with density values of 200, 150, 300, 100, and 250, respectively, its average density is (200 + 150 + 300 + 100 + 250) ÷ 5 = 200. The type distribution of data cells within the node is calculated (e.g., 40% text, 50% images, and 10% structured data). The type diversity is calculated (e.g., using Shannon entropy: -0.4ln0.4 -0.5ln0.5 -0.1ln0.1). Higher diversity indicates a greater cardinality in association strength. The classification level distribution is analyzed (e.g., 30% public, 50% internal, and 20% secret). The proportion of high-class classification (Secret and above) is calculated. The higher the proportion, the greater the association strength weight. The proportion of high-frequency data is calculated (e.g., 60% data with more than 100 calls). This proportion is positively correlated with association strength. The indicators of the three types of labels are normalized (such as type diversity ∈ [0, 2], high confidentiality ratio ∈ [0, 1], and high frequency ratio ∈ [0, 1]), and weight coefficients are set (such as type diversity weight 0.3, confidentiality weight 0.4, and frequency weight 0.3); the weighted summation is used to obtain the characteristic association strength value = 0.3×type entropy + 0.4×high confidentiality ratio + 0.3×high frequency ratio.
[0036] Calculation of node weight parameter β value: Normalize the node density value to the global maximum density, obtaining a normalized density value ∈ [0, 1]. For example, if the global maximum density is 500 and a node density is 200, the normalized value is 200 ÷ 500 = 0.4. Normalize the feature association strength value to the theoretical maximum value (for example, assuming the maximum strength is 2), obtaining a normalized association strength ∈ [0, 1].
[0037] Comprehensive calculation of β value: β=α×normalized density+(1-α)×normalized association strength, where α is the density weight coefficient (e.g., α=0.6). If a tree node has child nodes, the β value can be obtained by weighted average of the β values of the child nodes (e.g., weighted by the number of grids covered by the child nodes), ensuring that the β values of high-level nodes in the tree structure reflect the overall importance.
[0038] Each node records the following information: Spatial range (X / Y / Z range); Density value and normalized density; Feature association strength values and normalized strength; β value; List of child nodes.
[0039] Convert the property tree into JSON or XML format for easy storage and transmission.
[0040] By combining density and attribute correlation strength, the β value accurately reflects the business value and sensitivity of a set of data units. For example, data nodes with high density, high confidentiality, and high frequency have higher β values and are prioritized as core archiving objects. The attribute tree's hierarchical structure is dynamically generated based on data distribution. High-density areas are automatically subdivided into more subnodes (such as areas with high-frequency confidential data), implementing a differentiated strategy of "fine management of important data and extensive management of common data" to improve storage and retrieval efficiency. Attribute correlation strength calculation converts implicit associations between data type, confidentiality level, and frequency into explicit numerical values. For example, discovering the pattern of "high frequency association of image data with confidentiality level" provides a basis for optimizing data classification rules. The attribute tree presents data distribution logic in a visual hierarchy. Business personnel can quickly understand the source of data importance through node β values (for example, whether a node's high β value is due to high density or high confidentiality level). The tree structure can also be expanded by adding new tag dimensions (such as the "business type" tag).
[0041] In a preferred embodiment of the present invention, based on the β value of each node and the preset classification threshold, the grid coordinate mapping data unit set is subjected to differential processing to obtain a processed data unit set, which may include: Read the β value of each node in the attribute tree structure with β value, and divide the node level according to the preset threshold range, including: Define nodes with β value greater than or equal to the first preset threshold as high-β value nodes; Define nodes with β value greater than or equal to the second preset threshold and less than the first preset threshold as medium-β value nodes; Define nodes with β value less than the second preset threshold as low-β value nodes; Perform differential processing on the grid coordinate mapping data unit set, including enhancing the desensitization intensity of sensitive fields of data units associated with high-β value nodes; increasing the frequency of data cleaning execution for data units associated with medium-β value nodes; reducing the data sorting priority weight of data units associated with low-β value nodes, and output the optimized data unit set.
[0042] In an embodiment of the present invention, read the root node from the stored attribute tree structure (such as JSON or XML format), recursively traverse all child nodes, and obtain the β value, spatial coordinate range, and associated data unit set index of each node. For example, starting from the root node "X0-100_Y0-100_Z0-100", visit child nodes such as "X0-33_Y0-33_Z0-33" in sequence, and record the β value of each node (such as 0.8, 0.5, 0.2). Each tree node associates with a specific data unit set through the grid coordinate range. For example, the node "X66-100_Y66-100_Z66-100" corresponds to the last 34% area in the three-dimensional grid and associates with a list of all data unit IDs in this area (such as ID-001, ID-002...).
[0043] Node level division and threshold matching: The first preset threshold (Th1): such as 0.7 (threshold for high-β value nodes); The second preset threshold (Th2): such as 0.3 (threshold for dividing medium and low-β values).
[0044] High-β value nodes: β≥Th1 (such as β = 0.8≥0.7); Medium-β value nodes: Th2≤β<Th1 (such as 0.3≤β = 0.5<0.7); Low-β value nodes: β<Th2 (such as β = 0.2<0.3).
[0045] Enhance the desensitization intensity of high-β value node data units: Locate the fields that need to be desensitized according to the confidentiality level labels of data units (such as "secret", "confidential"). For example, ID numbers, bank account numbers, medical diagnosis records, etc. in government affairs data.
[0046] Upgrade the desensitization rules: Original desensitization rule: such as replacing the middle 8 digits of the ID number with " " (such as 110101 1234); Enhanced rules: The entire field is encrypted or replaced with an irreversible hash value (for example, "e10adc3949ba59abbe56e057f20f883e" encrypted using SHA-256).
[0047] Execution process: Traverse the data units associated with nodes with high β values (such as ID-001, ID-002); Parse the field structure of each data unit and identify sensitive fields; Perform desensitization according to the enhanced rules, overwriting the original desensitization results; Record desensitization logs (such as desensitization time and rule version).
[0048] The cleaning frequency of data units of medium beta nodes increases: The regular cleaning frequency is once a day, including filtering of invalid values (such as null values and abnormal characters) and format standardization (such as the date format is unified to YYYY-MM-DD).
[0049] Increase the frequency to 3 times a day: Added three scheduled cleaning tasks in the morning, afternoon and evening. Each cleaning task includes: Duplicate value detection (such as duplicate submission records of the same approval form); Cross-field logic validation (e.g. consistency check between the "age" field and the "date of birth" field); Semantic error correction (such as verification of the checksum of the "ID number" field).
[0050] Incremental cleaning strategy: Only data units updated since the last cleanup are cleaned, reducing resource consumption. For example, the third cleanup only processes data modified after 18:00 on the same day.
[0051] The sorting priority of the data units of nodes with low β values is reduced: Sort by call frequency in descending order, with high-frequency data displayed first (for example, the first 10 query results are high-frequency data).
[0052] Priority adjustment logic: The ranking weight coefficient α is adjusted: the original α = 1 (high-frequency data), and the α for low-β data is reduced to 0.5; the new priority = call frequency × α + other factors (such as update time). For example, low-β data A: with a call frequency of 100 times, the original ranking score is 100 × 1 = 100; the adjusted score is 100 × 0.5 = 50, and the ranking position moves from 5th to 15th.
[0053] Verification and integration of processed data unit sets: Randomly sample 10% of high-β data units to check whether the desensitization strength meets the enhanced rules (for example, whether the full-field encryption rate is 100%). Spot-check the cleaning logs of medium-β data units to confirm whether they are cleaned three times daily. Verify that the sorting results of low-β data meet the weight reduction requirements. Reorganize the processed data unit set according to the original grid coordinate structure to generate a new data set with processing marks (for example, add a "processing level" field to the data unit metadata).
[0054] Data on high-beta nodes meets the requirements of Level 3 of Information Security Protection 2.0 through enhanced masking (such as full-field encryption). For example, the masking intensity of "personal medical records" in government systems has been upgraded from partial replacement to irreversible encryption. Increased cleaning frequency for medium-beta data has reduced the data integrity error rate from 5 to below 1. For example, the format error rate of the "amount" field in corporate financial data has decreased. Low-beta data has been prioritized lower, increasing query response speed for high-frequency, high-value data (high beta) by 40% while reducing redundant computing resources for low-value data (e.g., CPU utilization by 15%). Dynamic processing strategies are adjusted based on beta, achieving an adaptive match between data importance and processing intensity. For example, quarterly financial report data (high beta) automatically triggers the highest level of masking, while historical transaction logs (low beta) retain only basic storage, reducing management costs.
[0055] In a preferred embodiment of the present invention, the processed data unit set is converted into a file set, and a differentiated blockchain evidence identification is added based on the attribute tree node β value, and a pre-archived data package with a β identification is output, which may include: Converting the optimized data unit set into a structured file set; Generate differentiated blockchain identifiers based on the attribute tree node β value, specifically including: Strong verification and evidence identification is added to the files corresponding to high β value nodes; Add standard evidence identification to the corresponding files of the medium β value nodes; Add basic evidence identification to the corresponding files of low β value nodes; Attach a corresponding identifier to each structured file and output a pre-archived data package with a β identifier.
[0056] In an embodiment of the present invention, the processed data unit set is traversed and the storage format is determined according to the data label (such as text, image, structured table): Structured data (such as database tables): Convert to XML or JSON format, preserving field structure and relationships. For example, when government approval process data is converted to XML, it is stored in the hierarchy of "approval node → person in charge → attachment"; Semi-structured data (such as log files): Parsed into a JSON object array, with key-value pairs retaining the original fields (such as "Timestamp: 2025-06-16" and "Operation Type: Query"); Unstructured data (such as images and audio): Keep the original binary format and attach metadata tags (such as file name, creation time, data type).
[0057] File set organization rules: The data is stored in folders according to grid coordinates. For example, the three-dimensional coordinate (15, 20, 5) corresponds to the folder "X15_Y20_Z5", which contains all data files in this area. The naming rule for each file is "grid coordinate_data unit ID_β value. format suffix", such as "X15_Y20_Z5_ID001_0.85.xml".
[0058] Access a consortium chain or private chain network (such as a government blockchain platform), obtain node communication permissions, and configure consensus algorithms (such as PBFT) and encryption certificates.
[0059] Set the evidence storage strategy corresponding to different β values: High β value (≥0.7): calling smart contracts to generate strong verification tokens; Medium β value (0.3-0.7): Generates standard evidence identification; Low β value (<0.3): Generates basic evidence identification.
[0060] Strong verification certificate identification generation (high β value node): Perform a SHA-256 hash on the file contents to obtain a fixed-length digest (e.g., a 64-bit hexadecimal string). Call the blockchain node API to obtain the current UTC time (e.g., "2025-06-16T08:30:00Z") and concatenate it with the hash value. Use both the file creator's private key and the blockchain node's private key for dual signing to generate a signature string. The identification structure is: {hash value, timestamp, signature string, evidence level: "strong verification", beta value: 0.85}.
[0061] Standard proof identification generation (medium β value node): Only SHA-256 hash + timestamp + single node signature is executed, and the identification structure is {hash value, timestamp, signature string, evidence level: "standard", β value: 0.5}.
[0062] Basic evidence identification generation (low β value node): Only the file hash value and blockchain block height are recorded. The identification structure is {hash value, block height, evidence level: "basic", β value: 0.2}.
[0063] Differentiated identification and data packet encapsulation: Structured / semi-structured files: Add an XML / JSON formatted evidence identification node to the file header, for example, insert it under the root node of an XML file. <blockchain-marker> ...< / blockchain-marker> Label; Unstructured files: Generates an independent marker file (.marker format) and stores it in the same path as the original file. For example, "image.jpg" corresponds to "image.jpg.marker".
[0064] Pre-archived data package encapsulation: According to the grid coordinate folder level, use ZIP or TAR format to package all files and identifiers; add a global index file (index.json) to the package file to record the mapping relationship of "grid coordinates-file list-β value-evidence level"; perform MD5 verification on the entire data package and generate a verification code attached to the package name (such as "archive_X0-33_Y0-33_Z0-33_20250616.md5").
[0065] Packet integrity and compliance verification: Parse each file's identifier and use a blockchain browser to check whether the hash value is on-chain (e.g., input the hash value and verify the consistency of the block height and timestamp). Check whether the signature string can be decrypted by the corresponding public key to ensure it has not been tampered with. Check whether the strong verification mark of high-beta files contains double signatures. Check whether the evidence mark of sensitive data (such as ID number) matches the redaction level (for example, fully encrypted data must have a corresponding strong verification mark).
[0066] Blockchain hashing ensures that any modification to high-beta data (such as core business records) results in a hash change, making it tamper-proof once uploaded to the blockchain. High-beta data consumes more computing power to perform strong validation (such as double signing), while low-beta data only records the basic hash, ensuring the robustness of the evidence for critical data. The evidence identifier includes elements such as a timestamp and signature, meeting the legal validity requirements for electronic archives. For example, during financial audits, the identifier can be used to quickly trace the creation and modification history of transaction data. Data packet index files are linked to beta values, enabling rapid filtering of high-value data by evidence level (strong validation / standard / basic). Both structured and unstructured data are converted into standardized file sets and affixed with blockchain identifiers, resolving the heterogeneity of native data formats and enabling archiving compatibility across systems and platforms.
[0067] In a preferred embodiment of the present invention, the data packets are pushed to the pre-archiving verification terminal, hierarchical verification is performed, the verified data packets are output, the passed data packets are parsed and imported into the database, storage resource priorities are allocated, and archived data with storage level identification is output, which may include: The pre-archived data package with the β mark is pushed to the pre-archived verification terminal. The verification process of the corresponding strength is triggered according to the β mark level. If the verification passes, the parsing and import are performed to obtain the parsed data; The parsed data is assigned storage resource priority based on the β identification value, imported into the storage database, and a storage level identification corresponding to the β value is attached to the archived data entering the database, and the archived data with the storage level identification is output.
[0068] In an embodiment of the present invention, a pre-archived data packet (such as a ZIP format) with a β identifier is pushed from the generation end to the pre-archived verification end server via FTP / SFTP or an API interface, and TLS encryption is enabled during the transmission process to ensure data integrity. For example, in a certain government data archiving system, high-β value data packets (such as social security sensitive data) are transmitted via a dedicated line, and low-β value data packets (such as public policy documents) are transmitted via a normal network. After receiving the data packet, the verification end first reads the global index file (index.json) and parses the β value and evidence level (strong verification / standard / basic) of each file. For example, β=0.85 is extracted from the file "X15_Y20_Z5_ID001_0.85.xml" and is determined to be high-β value node data.
[0069] Preset validation rule mapping:
[0070] Strong checksum execution details (high beta data): Compare the data packet with the originally generated checksum through the MD5 checksum to ensure that the transmission is not damaged; parse the blockchain evidence identification, call the blockchain node API to verify the hash value, timestamp and signature (such as using the public key to decrypt the signature string and compare whether the hash values are consistent); use an antivirus engine (such as Kaspersky) to scan the file for malicious code, and check the desensitization strength of sensitive fields (such as whether confidential data is fully encrypted); compare the data unit label with the actual content, such as whether the "ID number" field complies with the 18-digit rule, and whether the medical record contains the required diagnosis field.
[0071] Medium checksum execution details (medium beta data): Omit deep virus detection in security scans and perform only quick scans. Semantic checks are simplified to field format verification (for example, whether the date format is YYYY-MM-DD).
[0072] Weak check execution details (low beta data): Only verify the data packet MD5 and file format (such as whether there are tag closing errors in XML), and skip the deep verification of blockchain evidence.
[0073] After the verification process is completed, a verification report (including verification time, β value, and pass items) is generated, such as "High β value data packet ID-20250616-001 passed verification, taking 120 seconds"; a "Verification passed" timestamp is added to the data packet that passed the verification, marking it as importable.
[0074] Handling of verification failure: Low β value data: One automatic retry is allowed (such as retransmitting the data packet). If it still fails, it will be marked as "pending manual processing"; For data with medium or high beta values, the manual review process is immediately triggered. The verifier manually checks the cause of the error (for example, invalid signature may be due to private key expiration), corrects it, and resubmits the verification.
[0075] Data packet parsing and database import: For XML / JSON files, use a parser (such as a DOM parser) to extract data fields and map them to database table structures. For example, the "Applicant Name" field in the government approval XML file is mapped to the "applicant_name" database column. For unstructured files (such as images), thumbnails are generated and the paths are stored, and metadata (such as shooting time and resolution) is stored in database fields.
[0076] Batch import optimization: High-β data uses single-entry import + real-time index update to ensure immediate data availability; Medium and low beta value data is imported in batches (for example, a transaction is submitted every 1,000 records) to reduce database connection overhead.
[0077] Storage media mapping rules: High β value (β ≥ 0.7): Assigned to the SSD high-speed storage array, RAID level is RAID-10 (balancing speed and redundancy); Medium β value (0.3≤β<0.7): allocated to SAS hard disks, RAID-5; Low β value (β < 0.3): allocated to SATA hard disks or archival storage (such as tape libraries).
[0078] Storage class identifier generation: Add a "storage_level" field to the database table, with values of "high / medium / low" corresponding to the β value range. For example, β = 0.85 is marked as "high" and the storage path is " / ssd / storage / high / 2025 / 06 / "; At the same time, the index file is updated to record the storage location and level.
[0079] A storage management system (such as Nagios) monitors storage usage at each tier. When high-tier storage utilization exceeds 80%, an expansion alert is automatically triggered. Low-beta data is scanned regularly (e.g., monthly). If no access record has been maintained for six months, it is automatically migrated to offline archival storage, freeing up online storage resources. A robust validation process ensures the integrity and authenticity of highly sensitive data (such as bank transaction records). Combined with SSD-encrypted storage, this data meets the requirements of Level 4 of the Information Security Protection Technology 2.0. Adaptive validation and storage strategies are employed for medium- and low-sensitivity data to avoid excessive protection and resource waste. Storage tier identification is linked to beta values, allowing operations personnel to quickly locate core business data based on the "high storage tier," facilitating backup strategy development (e.g., daily incremental backups for high-tier data and weekly full backups for low-tier data). In the event of a failure, high-tier data is prioritized for recovery, minimizing business disruptions. The tiered storage architecture supports on-demand expansion of storage resources at different tiers (e.g., adding SSD arrays to accommodate high-beta data growth), avoiding the wasteful "full capacity expansion" of traditional unified storage architectures.
[0080] An embodiment of the present invention further provides a computing device comprising: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, executes the system described above. All implementations in the above system embodiments are applicable to this embodiment and can achieve the same technical effects.
[0081] The embodiment of the present invention further provides a computer-readable storage medium storing instructions, which, when executed on a computer, causes the computer to execute the system described above. All implementations in the above system embodiments are applicable to this embodiment and can achieve the same technical effects.
[0082] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A multimodal data intelligent archiving and management system, characterized in that: include: The data mapping module is used to extract multimodal data from the native business environment, generate a data unit set, construct a virtual reference plane, establish a reference rectangular grid topology structure and divide the dynamic coordinate axis in the plane, map the data unit set to grid coordinates, and output the grid coordinate mapped data unit set; The attribute tree construction module is used to map the density distribution of the data unit set according to the grid coordinates, construct a multidimensional attribute tree, generate dynamic evaluation nodes, calculate the weight parameter β value of each node, and output the attribute tree structure with β value; A differential processing module is used to perform differential processing on the grid coordinate mapping data unit set according to the β value of each node and the preset classification threshold to obtain a processed data unit set; The data packet generation module is used to convert the processed data unit set into a file set, add a differentiated blockchain evidence identification based on the attribute tree node β value, and output a pre-archived data packet with the β identification; The verification module is used to push data packets to the pre-archiving verification terminal, perform hierarchical verification, output the data packets that have passed the verification, parse the passed data packets and import them into the database, allocate storage resource priorities, and output archived data with storage level identification; The dynamic adjustment module is used to provide retrieval services for archived data, collect user access behavior data in real time, feed the access behavior data back to the attribute tree construction module, and dynamically adjust the coordinate axis division rules.
2. The multimodal data intelligent archiving and management system according to claim 1, characterized in that: Extract multimodal data from the native business environment, generate a data unit set, and construct a virtual reference plane. Establish a reference rectangular grid topology structure and divide the dynamic coordinate axis within the plane. Map the data unit set to grid coordinates and output the grid coordinate mapping data unit set, including: Extract multimodal data from the native business environment and generate independent data units with data type labels, confidentiality level labels, and call frequency labels; Perform preprocessing operations on independent data units, including screening valid data units based on preset rules, performing cleaning operations on invalid fields, sorting sensitive fields according to confidentiality requirements and data units according to call frequency, and outputting a set of preprocessed data units; Taking the preprocessed data unit set as input data, a three-dimensional virtual reference plane is constructed based on the data type label, the confidentiality level label and the frequency label; A benchmark rectangular grid topology structure covering the entire data range is established in the virtual reference plane, and the dynamic coordinate axis scale is divided according to the distribution density of the input data in the three-dimensional feature space. Each preprocessed data unit is mapped to the corresponding coordinate unit of the grid, and the grid coordinate mapping data unit set is output.
3. The multimodal data intelligent archiving and management system according to claim 2, characterized in that: A reference rectangular grid topology covering the entire data range is established in the virtual reference plane. The dynamic coordinate axis scale is divided according to the distribution density of the input data in the three-dimensional feature space. Each preprocessed data unit is mapped to the corresponding coordinate unit of the grid, and a grid coordinate mapping data unit set is output, including: Based on the three-dimensional space coordinate range of the pre-processed data unit set in terms of data type, confidentiality level, and call frequency, a rectangular grid structure that completely covers all data coordinates is generated; According to the distribution density of each preprocessed data unit in the three-dimensional space, the scale interval of each coordinate axis is adaptively adjusted, and each preprocessed data unit is positioned to the corresponding coordinate unit in the grid according to the data type label value, confidentiality level label value, and call frequency label value; Integrate all data units that have completed coordinate positioning to form a grid coordinate mapping data unit set.
4. The multimodal data intelligent archiving and management system according to claim 3, characterized in that: According to the density distribution of the grid coordinate mapping data unit set, a multidimensional attribute tree is constructed, dynamic evaluation nodes are generated, and the weight parameter β value of each node is calculated. The attribute tree structure with β value is output, including: Map the data unit set according to the grid coordinates, and count the number of data units in each grid unit as the density value; A multidimensional attribute tree is constructed based on the number of data units in each grid cell as the density value, where each tree node corresponds to a grid cell and the node level is determined by the position of the grid cell in the three-dimensional feature space; Generate a dynamic evaluation node for each tree node, obtain the density value of the grid unit corresponding to the corresponding node, and calculate the characteristic association strength value of the data unit in the corresponding grid unit based on the association relationship between the three types of labels: data type, confidentiality level and call frequency; According to the density value of the corresponding grid unit of the corresponding node and the characteristic correlation strength of the data unit in the corresponding grid unit, the node weight parameter β value is obtained, and a multidimensional attribute tree structure carrying the β value of each node is output.
5. The multimodal data intelligent archiving and management system according to claim 4, characterized in that: According to the β value of each node and the preset classification threshold, the grid coordinate mapping data unit set is subjected to differential processing to obtain the processed data unit set, including: Read the β value of each node in the attribute tree structure with β value, and divide the node level according to the preset threshold range, including: defining nodes whose β value is greater than or equal to a first preset threshold as high β value nodes; A node whose β value is greater than or equal to the second preset threshold and less than the first preset threshold is defined as a medium β value node; defining a node whose β value is less than a second preset threshold as a low β value node; Differentiated processing is performed on the grid coordinate mapping data unit set, including enhancing the desensitization strength of sensitive fields for data units associated with high β value nodes; increasing the frequency of data cleaning execution for data units associated with medium β value nodes; and reducing the data sorting priority weight for data units associated with low β value nodes, and outputting the optimized data unit set.
6. The multimodal data intelligent archiving and management system according to claim 5, characterized in that: Convert the processed data unit set into a file set, add a differentiated blockchain evidence identification based on the attribute tree node β value, and output a pre-archived data package with the β identification, including: Converting the optimized data unit set into a structured file set; Generate differentiated blockchain identifiers based on the attribute tree node β value, specifically including: Strong verification and evidence identification is added to the files corresponding to high β value nodes; Add standard evidence identification to the corresponding files of the medium β value nodes; Add basic evidence identification to the corresponding files of low β value nodes; Attach a corresponding identifier to each structured file and output a pre-archived data package with a β identifier.
7. The multimodal data intelligent archiving and management system according to claim 6, characterized in that: Push data packets to the pre-archiving verification terminal, perform hierarchical verification, output the verified data packets, parse and import the passed data packets into the database, allocate storage resource priorities, and output archived data with storage level identification, including: The pre-archived data package with the β mark is pushed to the pre-archived verification terminal. The verification process of the corresponding strength is triggered according to the β mark level. If the verification passes, the parsing and import are performed to obtain the parsed data; The parsed data is assigned storage resource priority based on the β identification value, imported into the storage database, and a storage level identification corresponding to the β value is attached to the archived data entering the database, and the archived data with the storage level identification is output.
8. A computing device, characterized in that include: one or more processors; A storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the system according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program, which, when executed by a processor, implements the system according to any one of claims 1 to 7.
Citation Information
Patent Citations
Data storage computing method and system
CN104731796A
Space meshing based government affair big data mining method
CN105279260A
Meteorological metadata storage method and system based on machine learning
CN120104579A
Emulating manual system of filing using electronic document and electronic file
WO2016060547A1