Information processing device, information processing method, and computer program

WO2026203002A1PCT designated stage Publication Date: 2026-10-01EXDATA INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/011400
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2026-10-01

Smart Images

  • Figure JP2025011400_01102026_PF_FP_ABST
    Figure JP2025011400_01102026_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device comprises a data acquisition unit, a data division unit, a scheme selection unit, a compression execution unit, and an aggregation unit. The data acquisition unit acquires structured data. The data division unit divides the structured data into a plurality of pieces of partial data. The scheme selection unit assigns a compression scheme to each piece of partial data on the basis of characteristics of each piece of partial data. The compression execution unit compresses each piece of partial data by using the assigned compression scheme. The aggregation unit generates one archive by aggregating the compressed pieces of partial data.
Need to check novelty before this filing date? Find Prior Art

Description

Information Processing Apparatus, Information Processing Method, and Computer Program

[0001] The technology disclosed in the present specification relates to information processing that handles structured data.

[0002] Structured Data is data having a predetermined structure, for example, data in CSV format or JSON format. One type of structured data is Real World Data. Real-world data is data recording various phenomena occurring in the real world, such as passenger flow data, logistics data, meteorological data, power data, and the like. Real-world data is also referred to as Spatio-Temporal Data. In many industrial fields, real-world data is collected for some purpose such as marketing, congestion relief, crime prevention, and power distribution optimization, for example.

[0003] When collecting structured data such as real-world data, data compression is performed to reduce data storage costs and the like (see, for example, Patent Document 1).

[0004] Japanese Patent Application Laid-Open No. 2011-090526

[0005] Conventionally, when compressing structured data, a single compression method is used, so there was room for improvement in compression rate.

[0006] The present specification discloses a technology capable of solving the above-mentioned problems.

[0007] The technology disclosed in the present specification can be implemented, for example, in the following forms.

[0008] (1) An information processing apparatus disclosed in the present specification includes a data acquisition unit, a data division unit, a method selection unit, a compression execution unit, and an aggregation unit. The data acquisition unit acquires structured data. The data division unit divides the structured data into a plurality of partial data. The method selection unit allocates a compression method to each of the partial data based on characteristics of each of the partial data. The compression execution unit compresses each of the partial data using the allocated compression method. The aggregation unit collects the compressed partial data together to generate one archive.

[0009] According to this information processing device, compression is performed on each of the multiple sub-data points contained in the structured data using a compression method assigned based on the characteristics of the sub-data point, and the compressed sub-data points are combined into a single archive. Therefore, compared to a configuration in which compression is performed on the entire structured data using a single compression method, it is possible to achieve a higher compression ratio for the structured data.

[0010] (2) In the above-mentioned information processing device, the characteristics may include at least one of the data type and the data arrangement. With this configuration, an appropriate compression method can be assigned to each partial data based on at least one of the partial data type and the arrangement, thereby enabling compression of structured data at a higher compression ratio.

[0011] (3) In the above-mentioned information processing device, the data acquisition unit may acquire metadata that identifies the characteristics of each of the subdata included in the structured data, and the method selection unit may assign the compression method based on the metadata. With this configuration, the characteristics of the subdata can be grasped quickly and reliably, and an appropriate compression method can be assigned to each subdata, thereby achieving compression of structured data at a higher compression ratio.

[0012] (4) In the above-mentioned information processing device, the aggregation unit may generate the archive including the metadata. With this configuration, it is possible to generate an archive that includes each portion of data compressed with a high compression ratio and also includes metadata, and an archive that can be used for data utilization such as data conversion, analysis, and visualization can be obtained.

[0013] (5) In the above-mentioned information processing device, the data acquisition unit may generate the metadata based on the structured data. With this configuration, it is possible to achieve high compression ratios even for structured data that does not yet have metadata attached.

[0014] (6) In the above-mentioned information processing device, the data acquisition unit may analyze the structured data to create structural metadata from the metadata of the structured data, create semantic metadata from the metadata of the structured data based on the structural metadata, and merge the structural metadata and the semantic metadata to generate the metadata. With this configuration, metadata can be attached with high accuracy based on the structured data.

[0015] (7) The above-mentioned information processing device may further include a storage unit that stores correspondence relationship information indicating the correspondence relationship between the characteristics and the compression method, and the method selection unit may refer to the correspondence relationship information to assign the compression method. With this configuration, an appropriate compression method can be assigned to each part of the data quickly and reliably, and structured data can be compressed at a high compression ratio.

[0016] (8) The above-mentioned information processing device may further include an information update unit for updating the correspondence relationship information. With this configuration, as new compression methods are developed, a more appropriate compression method can be assigned to each part of the data, and a higher compression ratio for structured data can be achieved.

[0017] (9) In the above-mentioned information processing device, the correspondence information may be information that shows the correspondence between the characteristics, the compression method, and an index value that correlates with the compression ratio. With this configuration, an appropriate compression method can be quickly and reliably assigned to each part of the data according to the required data accuracy, and structured data can be compressed at a high compression ratio.

[0018] (10) The above-mentioned information processing device may further include an unpacking execution unit that unpacks the archive to obtain the structured data before compression. This configuration makes it possible to unpack structured data that has been compressed with a high compression ratio.

[0019] The technologies disclosed herein can be implemented in various forms, for example, as information processing apparatus and method, data compression apparatus and method, data decompression apparatus and method, data conversion apparatus and method, data analysis apparatus and method, data visualization apparatus and method, metadata generation apparatus and method, computer programs that implement the above methods, and non-temporary recording media on which the computer programs are stored.

[0020] This embodiment shows an overview of the structured data compression method. It also shows a schematic diagram of the information processing device 100. The diagrams illustrate the metadata generation process, the flowchart, the overview of the metadata generation process, the compression process, the compression process, an example of correspondence relationship information (CR), an example of correspondence relationship information (CR), the unfolding process, the evaluation results for the generation accuracy of semantic metadata, the evaluation results for the data compression ratio using simulation data, the evaluation results for the data compression ratio using simulation data, the evaluation results for the data compression ratio using simulation data, the evaluation results for the data compression ratio using simulation data, the evaluation results for the data compression ratio using actual data, the evaluation results for the data compression ratio using actual data, the evaluation results for the average compression ratio when using actual data, the evaluation results for the compression speed targeting simulation data, the evaluation results for the average compression speed targeting actual data, the evaluation results for the unfolding speed targeting simulation data, and the evaluation results for the average unfolding speed targeting actual data.

[0021] (Outline of Structured Data Compression Method) Figure 1 is an explanatory diagram showing an overview of the structured data compression method in this embodiment. Structured data is data that has a predetermined structure, such as data in CSV format or JSON format. Structured data may include portions in which data of the same type and class appears repeatedly a predetermined number of times or more (for example, 100 times or more). Structured data may be data of a predetermined size or larger (for example, 10 MB or more) in its original state (uncompressed state). In this embodiment, real-world data DRW will be used as an example of structured data. Real-world data DRW is data that records various events that occur in the real world.

[0022] As shown in Figure 1, in this embodiment, metadata Prw is automatically generated from real-world data DRW. Metadata Prw is data that describes information about the data, for example, it indicates the characteristics of each subdata (field, property, column, etc.) contained in the real-world data DRW. Examples of data characteristics include the data name, precision, numerical range, unit, type, order, scale, and a detailed description of the data. Examples of data types include integers, floating-point numbers, strings, classes / lists, Booleans, and datetimes. Examples of data orders include constant, sequential, random, and enumerated. Order: Constant may be "approximately" constant. Sorting: Sequential data can also be classified by its degree of sequentiality. For example, it can be divided into sequential data of 75% or more (where the difference between the values ​​before and after 75% is constant) and sequential data of less than 75% (where the difference between the values ​​before and after less than 75% is constant).

[0023] In this embodiment, compression of the real-world data DRW is performed by referring to the metadata PRW. First, based on the metadata PRW, the real-world data DRW is divided into multiple subdata. In the example in Figure 1, the real-world data DRW is divided into subdata da of the data field "aaa", subdata db of the data field "bbb", subdata dc of the data field "ccc", and the remaining subdata dr (non-repeating data field). Subdata da is of type: integer value, arranged sequentially. Subdata db is of type: integer value, arranged constant. Subdata dc is of type: string data.

[0024] Based on the characteristics of each subdata segment identified by metadata Prw, a compression method is assigned to each subdata segment, and compression is performed using the assigned method. In the example in Figure 1, subdata segment da is assigned compression method A (a method suitable for sequential numerical data), subdata segment db is assigned compression method B (a method suitable for constant numerical data), subdata segment dc is assigned compression method C (a method suitable for string data), and the remaining subdata segment dr is assigned a common compression method. The compression process for each subdata segment generates compressed subdata segments da(c), db(c), dc(c), and dr(c). The compression process for each subdata segment can be executed in parallel.

[0025] Finally, a single archive, Dar, is generated by combining the compressed partial data da(c), db(c), dc(c), and dr(c). In this embodiment, metadata Prw is included in the archive Dar.

[0026] As described above, in this embodiment, compression processing is performed on each of the multiple partial data contained in the real-world data DRW using a compression method assigned based on the characteristics of the partial data, and the compressed partial data is combined into a single archive Dar. Therefore, compared to an embodiment in which compression processing is performed on the entire real-world data using a single compression method, it is possible to achieve a higher compression ratio for the real-world data. The embodiment will be described in more detail below.

[0027] (Configuration of Information Processing Device 100) Figure 2 is an explanatory diagram showing the schematic configuration of the information processing device 100. The information processing device 100 is composed of, for example, a computer (PC, server, smartphone, tablet terminal, etc.).

[0028] The information processing device 100 comprises a control unit 110, a storage unit 120, a display unit 130, an operation input unit 140, and an interface unit 150. Each of these units is connected to the others via a bus 190 so as to be able to communicate with each other.

[0029] The display unit 130 of the information processing device 100 is composed of, for example, a liquid crystal display or an organic EL display, and displays various images and information. The display unit 130 is an example of an output device and display device. The operation input unit 140 is composed of, for example, a keyboard, mouse, buttons, microphone, trackpad, etc., and receives user operations and instructions. The display unit 130 may also function as the operation input unit 140 by being equipped with a touch panel. The interface unit 150 is composed of, for example, a LAN interface or a USB interface, and communicates with other devices by wired or wireless connection. The information processing device 100 may also be equipped with other output devices (for example, a speaker).

[0030] The storage unit 120 of the information processing device 100 is composed of, for example, ROM, RAM, a hard disk drive (HDD), and stores various programs and data, and is used as a working area and temporary data storage area when executing various programs. For example, the storage unit 120 stores correspondence relationship information CR, which is referenced when assigning compression methods to each part data contained in real-world data DRW. The storage unit 120 also stores data processing program CP, which is a computer program for executing various processes described later. The data processing program CP is provided, for example, stored on a computer-readable recording medium (not shown) such as a CD-ROM, DVD-ROM, or USB memory, or is provided in a state that can be obtained from an external device (a server on a network or other terminal device) via the interface unit 150, and is stored in the storage unit 120 in a state that can be operated on the information processing device 100.

[0031] The control unit 110 of the information processing device 100 is configured, for example, with a CPU, and controls the operation of the information processing device 100 by executing a computer program read from the storage unit 120. For example, the control unit 110 functions as a data processing unit 111 for executing various processes described later by reading and executing a data processing program CP from the storage unit 120. The data processing unit 111 includes a data acquisition unit 112, a data splitting unit 113, a method selection unit 114, a compression execution unit 115, an aggregation unit 116, an information update unit 117, and an expansion execution unit 118. The functions of each of these units will be explained in accordance with the descriptions of various processes described later.

[0032] (Metadata Generation Process) Next, the metadata generation process performed by the information processing device 100 of this embodiment will be described. Figure 3 is a flowchart of the metadata generation process. Figure 4 is an explanatory diagram showing an overview of the metadata generation process. The metadata generation process is the process of generating metadata Prw for each subdata contained in the real-world data DRW, which is structured data. The metadata generation process is started when a user operates the operation input unit 140 of the information processing device 100 and inputs a start command.

[0033] First, the data acquisition unit 112 (Figure 2) of the information processing device 100 acquires the real-world data DRW to be processed (S110). The data acquisition unit 112 may acquire the real-world data DRW from an external source via the interface unit 150. The data acquisition unit 112 may also acquire at least a portion of the real-world data DRW by generating it itself. The acquired real-world data DRW is stored, for example, in the storage unit 120.

[0034] Next, the data acquisition unit 112 analyzes the real-world data Drw and creates structural metadata Pst from the metadata of the real-world data Drw (S120). Structural metadata Pst is information that identifies, for example, the root structure of the entire real-world data Drw, as well as the data type and arrangement of each field.

[0035] Next, the data acquisition unit 112 creates semantic metadata Pse from the metadata of the real-world data Drw (S130). Semantic metadata Pse is information that identifies, for example, the description of each field of the real-world data Drw, the units of numbers, the data format (for example, the date and time display format), etc. The data acquisition unit 112 obtains semantic metadata Pse by inputting, for example, the structural metadata Pst created in S120 and some sample data Dsmp sampled from the real-world data Drw into the LLM (Large Language Model) along with prompts.

[0036] Next, the data acquisition unit 112 merges the structural metadata Pst and the semantic metadata Pse to generate metadata Prw for the real-world data Drw (S140).

[0037] In this embodiment, the metadata Prw of the real-world data Drw is implemented as an extension of the Dataset metadata representation of schema.org. However, the metadata Prw may also be implemented as an extension of other known metadata representations such as SensorML's SWE, or as an original metadata representation that is not an extension of other metadata representations.

[0038] In this embodiment, in the schema, the class RealWorldDataFieldProfile represents the metadata Prw for each real-world data Drw. The property dbp:structure contains the class dbp:RealWorldDataStructureGraph. Other properties contain information such as the detailed format of JSON and CSV files. The class dbp:RealWorldDataStructureGraph describes a JSON-LD compliant data structure. The property @graph of the class dbp:RealWorldDataStructureGraph contains a list of classes (objects) that represent the data structure of real-world data. Each class (object) contained in @graph has the properties shown in Table 1 below. In this specification, including Table 1, "String" means string, "Class" means class, "List" means list, "Boolean" means boolean value, "Integer" means integer, "Structural" means structured, and "Semantic" means semantic.

[0039]

[0040] The following are specific examples of metadata PRW for real-world data DRW. The first example concerns the following CSV data as real-world data DRW. 2023-01-01T00:00:00, 6909fcbb-ef24-4f9b-8981-4a81297cb133, 100, 35.987903, 136.88374 2023-01-01T00:01:02, ba3364b5-662e-4909-b6d1-0e59d8b34beb, 101, 35.987008, 136.88371 2023-01-01T00:01:58, 72b81471-ed17-4f2d-8df7-1249f7daf08f, 102, 35.988905, 136.83374 2023-01-01T00:03:00, 2159cb54-b22d-4fa8-8a12-7eeeb41f4348, 103, 35.967909, 136.88344 2023-01-01T00:04:01, 98f1b873-31fc-4d6c-a3c6-3ca562830c14, 104, 35.917918, 136.85375

[0041] An example of representing metadata Prw for the above CSV-formatted real-world data Drw in JSON-LD format is shown below.

[0042] { "@id": "https: / / path.to / this.jsonld", "@type": "dbp:RealWorldDataFieldProfile", "@context": { "dbp": "https: / / exdata.co.jp / dbp / schema / ", "rdf": "https: / / www.w3.org / 1999 / 02 / 22-rdf-syntax-ns#", "rdfs": "http: / / www.w3.org / 2000 / 01 / rdf-schema#", "schema": "https: / / schema.org / " }, "schema:name": "Sample RWD Profile for CSV", "dbp:structure": { "@type": "dbp:RealWorldDataStructureGraph", "@graph": [ { "@id": "root", "@type": "rdf:List", "rdfs:label": "root", "schema:rangeIncludes": { "@id": "class1" }, "dbp:dbpaCompress": true, "dbp:dbpaCompressListID": 0 }, { "@id": "class1", "@type": "rdfs:Class", "rdfs:label": "class1", "schema:rangeIncludes": [ { "@id": "datetime" }, { "@id": "id" }, { "@id": "seq" },{ "@id": "latitude" }, { "@id": "longitude" } ], "schema:domainIncludes": { "@id": "root" }, "dbp:dbpaCompressRowID": 0 }, { "@id": "datetime", "@type": "dbp:RealWorldDataStructureProperty", "rdfs:label": "datetime", "dbp:itemType": "DateTime", "rdfs:comment": "Timestamp when an item was recorded.", "schema:rangeIncludes": { "@id": "schema:Date" }, "schema:domainIncludes": { "@id": "class1" }, "dbp:VariableCharacteristicEnumeration": "Interval", "dbp:dbpaDateTimeFormat": "%Y-%m-%dT%H:%M:%S", "dbp:dbpaCompressParentListID": 0, "dbp:dbpaCompressColumnID": 0 }, { "@id": "id", "@type": "dbp:RealWorldDataStructureProperty", "rdfs:label": "id", "dbp:itemType": "String(32)","rdfs:comment": "Unique identifier for each item.", "schema:rangeIncludes": { "@id": "schema:Text" }, "schema:domainIncludes": { "@id": "class1" }, "dbp:VariableCharacteristicEnumeration": "Nominal", "dbp:isEnumValue": false, "dbp:dbpaCompressParentListID": 0, "dbp:dbpaCompressColumnID": 1 }, { "@id": "seq", "@type": "dbp:RealWorldDataStructureProperty", "rdfs:label": "seq", "dbp:itemType": "Integer", "rdfs:comment": "Sequential Number of data record.", "schema:rangeIncludes": { "@id": "schema:Integer" }, "schema:domainIncludes": { "@id": "class1" }, "dbp:VariableCharacteristicEnumeration": "Ordinal", "dbp:dbpaCompressParentListID": 0, "dbp:dbpaCompressColumnID": 2, "dbp:RangeMin": "100", "dbp:RangeMax": "999999", "dbp:BaseIncrement": 1,"dbp:UseBaseIncrement": true }, { "@id": "latitude", "@type": "dbp:RealWorldDataStructureProperty", "rdfs:label": "latitude", "dbp:itemType": "Float", "rdfs:comment": "Geographic coordinate representing north-south position", "schema:unitText": "DD", "schema:rangeIncludes": { "@id": "schema:Float" }, "schema:domainIncludes": { "@id": "class1" }, "dbp:VariableCharacteristicEnumeration": "Proportional", "dbp:dbpaCompressParentListID": 0, "dbp:dbpaCompressColumnID": 3, "dbp:RangeMin": "35.15300000", "dbp:RangeMax": "35.15999999", "dbp:DecimalPlaces:": 6, "dbp:PrecisionBytes": 2 }, { "@id": "longitude", "@type": "dbp:RealWorldDataStructureProperty", "rdfs:label": "longitude", "dbp:itemType": "Float","rdfs:comment": "Geographic coordinate representing east-west position", "schema:unitText": "DD", "schema:rangeIncludes": { "@id": "schema:Float" }, "schema:domainIncludes": { "@id": "class1" }, "dbp:VariableCharacteristicEnumeration": "Interval", "dbp:dbpaCompressParentListID": 0, "dbp:dbpaCompressColumnID": 4, "dbp:RangeMin": "136.9700000", "dbp:RangeMax": "136.9899999", "dbp:DecimalPlaces:": 5, "dbp:PrecisionBytes": 2 } ] }, "schema:dateCreated": "2024-12-31T23:59:59.999999+09:00", "schema:encodingFormat": "text / csv", "dbp:csvHasHeader": false, "dbp:newLineCharacter": "LF"},

[0043] A second concrete example of metadata PRW for real-world data DRW is the following JSON data. [ { "dateStart": "2023-01-01T00:00:00", "dateEnd": "2023-01-01T01:59:59", "id": "6909fcbb-ef24-4f9b-8981-4a81297cb133", "deviceType": "ABC", "positions": [ { "latitude": 35.9879087, "longitude": 136.883749}, { "latitude": 35.9879086, "longitude": 136.883751}, ... ]}, { "dateStart": "2023-01-01T00:01:02", "dateEnd": "2023-01-01T02:01:01", "id": "ba3364b5-662e-4909-b6d1-0e59d8b34beb", "deviceType": "DEF", "positions": [ { "latitude": 35.9879035, "longitude": 136.883732}, { "latitude": 35.9879037, "longitude": 136.883730}, ... ]}, ... ]

[0044] An example of representing metadata Prw for the above JSON-formatted real-world data Drw in JSON-LD format is shown below.

[0045] { "@id": "https: / / path.to / this.jsonld", "@type": "dbp:RealWorldDataFieldProfile", "@context": { "dbp": "https: / / exdata.co.jp / dbp / schema / ", "rdf": "https: / / www.w3.org / 1999 / 02 / 22-rdf-syntax-ns#", "rdfs": "http: / / www.w3.org / 2000 / 01 / rdf-schema#", "schema": "https: / / schema.org / " }, "schema:name": "Sample RWD Profile for JSON", "dbp:structure": { "@type": "dbp:RealWorldDataStructureGraph", "@graph": [ { "@id": "root", "@type": "rdf:List", "rdfs:label": "root", "schema:rangeIncludes": { "@id": "class1" }, "dbp:dbpaCompress": true, "dbp:dbpaCompressListID": 0 }, { "@id": "class1", "@type": "rdfs:Class", "rdfs:label": "class1", "schema:rangeIncludes": [ { "@id": "dateStart" }, { "@id": "dateEnd" }, { "@id": "id" },{ "@id": "deviceType" }, { "@id": "points" } ], "schema:domainIncludes": { "@id": "root" }, "dbp:dbpaCompressRowID": 0 }, { "@id": "dateStart", "@type": "dbp:RealWorldDataStructureProperty", "rdfs:label": "dateStart", "dbp:itemType": "DateTime", "rdfs:comment": "Timestamp when an item was started to record.", "schema:rangeIncludes": { "@id": "schema:Date" }, "schema:domainIncludes": { "@id": "class1" }, "dbp:VariableCharacteristicEnumeration": "Interval", "dbp:dbpaDateTimeFormat": "%Y-%m-%dT%H:%M:%S", "dbp:dbpaCompressParentListID": 0, "dbp:dbpaCompressColumnID": 0 }, { "@id": "dateEnd", "@type": "dbp:RealWorldDataStructureProperty", "rdfs:label": "dateEnd", "dbp:itemType": "DateTime","rdfs:comment": "Timestamp when an item was ended to record.", "schema:rangeIncludes": { "@id": "schema:Date" }, "schema:domainIncludes": { "@id": "class1" }, "dbp:VariableCharacteristicEnumeration": "Interval", "dbp:dbpaDateTimeFormat": "%Y-%m-%dT%H:%M:%S", "dbp:dbpaCompressParentListID": 0, "dbp:dbpaCompressColumnID": 1 }, { "@id": "id", "@type": "dbp:RealWorldDataStructureProperty", "rdfs:label": "id", "dbp:itemType": "String(32)", "rdfs:comment": "Unique identifier for each item.", "schema:rangeIncludes": { "@id": "schema:Text" }, "schema:domainIncludes": { "@id": "class1" }, "dbp:VariableCharacteristicEnumeration": "Nominal", "dbp:isEnumValue": false, "dbp:dbpaCompressParentListID": 0, "dbp:dbpaCompressColumnID": 2 },{ "@id": "deviceType", "@type": "dbp:RealWorldDataStructureProperty", "rdfs:label": "deviceType", "dbp:itemType": "String(3)", "rdfs:comment": "Type of the device.", "schema:rangeIncludes": { "@id": "schema:Text" }, "schema:domainIncludes": { "@id": "class1" }, "dbp:VariableCharacteristicEnumeration": "Nominal", "dbp:isEnumValue": true, "dbp:dbpaCompressParentListID": 0, "dbp:dbpaCompressColumnID": 3 }, { "@id": "points", "@type": "dbp:RealWorldDataStructureProperty", "rdfs:label": "points", "dbp:itemType": "list", "rdfs:comment": "Location history of a record.", "schema:rangeIncludes": { "@id": "list2" }, "schema:domainIncludes": { "@id": "class1" }, "dbp:dbpaCompressParentListID": 0, "dbp:dbpaCompressColumnID": 4,"dbp:dbpaChildrenLists": [ 1 ] }, { "@id": "list2", "@type": "rdf:List", "rdfs:label": "list2", "schema:domainIncludes": { "@id": "points" }, "schema:rangeIncludes": { "@id": "class2" }, "dbp:dbpaCompress": true, "dbp:dbpaCompressListID": 1 }, { "@id": "class2", "@type": "rdfs:Class", "rdfs:label": "class2", "schema:rangeIncludes": [ { "@id": "latitude" }, { "@id": "longitude" } ], "schema:domainIncludes": { "@id": "list2" }, "dbp:dbpaCompressRowID": 1 }, { "@id": "latitude", "@type": "dbp:RealWorldDataStructureProperty", "rdfs:label": "latitude", "dbp:itemType": "Float", "rdfs:comment": "Geographic coordinate representing north-south position","schema:unitText": "DD", "schema:rangeIncludes": { "@id": "schema:Float" }, "schema:domainIncludes": { "@id": "class2" }, "dbp:VariableCharacteristicEnumeration": "Proportional", "dbp:dbpaCompressParentListID": 1, "dbp:dbpaCompressColumnID": 0, "dbp:RangeMin": "35.153000000", "dbp:RangeMax": "35.159999999", "dbp:DecimalPlaces:": 6, "dbp:PrecisionBytes": 2 }, { "@id": "longitude", "@type": "dbp:RealWorldDataStructureProperty", "rdfs:label": "longitude", "dbp:itemType": "Float", "rdfs:comment": "Geographic coordinate representing east-west position", "schema:unitText": "DD", "schema:rangeIncludes": { "@id": "schema:Float" }, "schema:domainIncludes": { "@id": "class2" }, "dbp:VariableCharacteristicEnumeration": "Interval","dbp:dbpaCompressParentListID": 1, "dbp:dbpaCompressColumnID": 1, "dbp:RangeMin": "136.97000000", "dbp:RangeMax": "136.98999999", "dbp:DecimalPlaces:": 7, "dbp:PrecisionBytes": 2 } ] }, "schema:dateCreated": "2024-12-31T23:59:59.999999+09:00", "schema:encodingFormat": "application / json"},

[0046] (Compression Processing) Next, the compression processing executed by the information processing apparatus 100 of the present embodiment will be described. FIG. 5 is a flowchart showing the compression processing. FIG. 6 is an explanatory diagram showing an outline of the compression processing. The compression processing is processing for compressing real-world data Drw with reference to metadata Prw. The compression processing is started in response to a user operating the operation input unit 140 of the information processing apparatus 100 and inputting a start instruction.

[0047] First, the data acquisition unit 112 (FIG. 2) of the information processing apparatus 100 acquires real-world data Drw to be processed and its metadata Prw (S210). The data acquisition unit 112 may acquire the real-world data Drw and the metadata Prw from the outside via the interface unit 150. The data acquisition unit 112 may acquire the real-world data Drw and the metadata Prw by generating at least a part thereof by itself. For example, after acquiring the real-world data Drw from the outside via the interface unit 150, the data acquisition unit 112 may acquire the data by generating the metadata Prw of the real-world data Drw by itself. The acquired real-world data Drw and metadata Prw are stored in the storage unit 120, for example.

[0048] Next, the data division unit 113 (Figure 2) of the information processing device 100 divides the real-world data Drw into multiple sub-data based on the metadata Prw (S220). In the example in Figure 6, the real-world data Drw is divided into multiple sub-data, including sub-data d1 of the "id" field, sub-data d2 of the "name" field, and sub-data d3 of the "temp" field. The non-list data is represented as the remaining sub-data dr.

[0049] Next, the method selection unit 114 (Figure 2) of the information processing device 100 assigns a compression method suitable for compressing each portion of data based on the characteristics of each portion of data (S230). The method selection unit 114 assigns the compression method by referring to the correspondence relationship information CR stored in the storage unit 120. The correspondence relationship information CR is information that shows the correspondence between data characteristics, compression methods, and index values ​​that correlate with the compression ratio.

[0050] Figures 7 and 8 are explanatory diagrams showing examples of correspondence information CR. The correspondence information CR (CR1) shown in Figure 7 relates to a lossless compression method, and the correspondence information CR (CR2) shown in Figure 8 relates to a lossy compression method. In the correspondence information CRs shown in Figures 7 and 8, the compression method to be assigned is specified for each combination of data type (integer value, floating point, string, date and time) and data arrangement (almost constant, 75% or more sequential, less than 75% sequential, random, enumeration).

[0051] The compression methods illustrated in Figures 7 and 8 are the following eight: (1) Middle-out Compression - A method in which the XOR of the new value with the previous value is calculated and the result of that calculation is saved. - If the value does not change, only 0 is saved, making it effective for data whose values ​​hardly change. - A method used in Facebook® Gorilla, applied at the byte level instead of the bit level.

[0052] (2) Irregular Delta Record - A method that saves the index of the part where the difference between the preceding and succeeding values ​​differs from the Base Increase defined in the metadata, and the difference between that part and the previous value. Effective for continuous value properties (fields) where the value increases (or decreases) at a nearly constant pace.

[0053] (3) Flexible Bytes Differential Encoding - A method that records the difference in numerical values. - Usable when the range of numerical variation falls within the range of an integer that can be represented by the number of bytes in DiffNumBites. - Compression ratio is higher for properties (fields) where DiffNumBites can be made small. In other words, it is effective for properties (fields) where the overall value is continuous, but the difference is often different each time.

[0054] (4) Range Limited Encoding - During encoding, calculate the value y = x - RangeMin by subtracting RangeMin from the original value x, and encode it as an integer of the calculated number of bytes. - During decoding, convert the encoded value into an integer of 32 bits to 128 bits from the values ​​of RangeMin and RangeMax, add RangeMin, and decode it back to the original value x = y + RangeMin.

[0055] (5) Range Precision Limited Encoding - An improved version of Range Limited Encoding for Integer, adapted for Float - Fixes the number of bytes after encoding to PrecisionBytes - During encoding, first calculate the value y = x - RangeMin by subtracting RangeMin from the original value x. Then, for each bit after encoding, when y is expressed in sum form, (RangeMax - RangeMin) / 2 1 , (RangeMax-RangeMin) / 2 2 , (RangeMax-RangeMin) / 2 3A bit of 1 or 0 is recorded depending on whether the value of , ... is included. During decoding, y (or an approximate value of it) is decoded from the encoded bit sequence, and the original value (or an approximate value of it) is obtained using x = y + RangeMin.

[0056] (6) Decimal Places Limited Encoding - This method encodes by specifying the number of decimal places to preserve using DecimalPlaces. - This method is effective when there are few decimal places or when it is necessary to strictly preserve the original value (lossless). - The number of bytes during encoding is determined by calculating the required number of integer bytes from the values ​​of RangeMin, RangeMax, and DecimalPlaces.

[0057] (7) Per-Value Dictionary Mapping - First, extract the words to be registered in the dictionary as all possible value patterns within the property (field) sequence. - Next, encode the number of value patterns as the length of the dictionary, then encode the length of the 0th word and the 0th word itself, and encode this up to the 1st, 2nd, ..., n-1th word. - Finally, encode the order in which the values ​​appear within the original property (field) sequence as the word index.

[0058] (8) Flexible Precision Timestamping - When the date and time string representations are consistent within a series of properties (fields), they are all converted to integer timestamp values ​​to reduce the number of bytes. Furthermore, the timestamp is saved as a difference from the initial value to further reduce the number of bytes. - During encoding, timestamping is done in the smallest unit specified by DatetimePrecision, and the difference is saved in the number of bytes specified by DatetimeDiffBytes. DatetimePrecision and DatetimeDiffBytes are values ​​specified when metadata is generated.

[0059] Furthermore, as shown in Figures 7 and 8 with "None," when the data type is an "integer value" and the arrangement is "random" or "enumerated," if the range of integers cannot be determined, or if the value itself is very small (1-2 bytes even when represented as a string), each value may be treated as a simple string, and no compression method may be assigned.

[0060] Similarly, as shown in Figures 7 and 8 with "None," when the data type is "Float," if the range of floating-point numbers cannot be determined or if the number of digits in the value varies frequently, Dicimal Places Limited Encoding may not be applicable, or applying it may actually increase the size. In such cases, it may be better to treat it as a string and not assign a compression method.

[0061] Similarly, as indicated by "None" in Figures 7 and 8, if the data type is "string" and the order is "approximately constant" or "random," then no compression method may be assigned.

[0062] Similarly, as indicated by "None" in Figures 7 and 8, if the data type is "Date and Time" and it is better to treat it as a string, then no compression method may be assigned.

[0063] As shown in the examples of correspondence information CR in Figures 7 and 8, the compression method (lossless or lossy method) to be assigned is specified for each combination of data type and arrangement. The method selection unit 114 refers to the correspondence information CR and assigns the compression method associated with the characteristics of each partial data to the corresponding data.

[0064] In this embodiment, the information update unit 117 (Figure 2) of the information processing device 100 can update the correspondence relationship information CR. For example, if a more suitable compression method is developed as a compression method to associate with the characteristics of each partial data, the information update unit 117 can update the correspondence relationship information CR to associate the compression method with the characteristics of the data.

[0065] Next, the compression execution unit 115 (Figure 2) of the information processing device 100 performs compression processing on each subdata using the assigned compression method (S240). In the example in Figure 6, compression method A (a method suitable for sequential numerical data) is assigned to subdata d1, compression method B (a method suitable for strings) is assigned to subdata d2, and compression method C (a method suitable for floating-point data) is assigned to subdata d3. The compression processing for each subdata can be executed in parallel. In the example in Figure 6, no compression method is assigned to subdata dr. Alternatively, after the compression processing using the assigned method in S230 (application of an intermediate compression layer), the compression execution unit 115 may uniformly perform a general compression process, such as xz compression, on each subdata.

[0066] Finally, the aggregation unit 116 (Figure 2) of the information processing device 100 combines the partial data after the compression process in S240, and adds metadata Prw to generate a single archive Dar (S250). As a result, an archive Dar is obtained in which each of the multiple partial data contained in the real-world data Drw has been compressed using an appropriate compression method in consideration of the characteristics of each partial data. Therefore, compared to a configuration in which the compression process is performed on the entire real-world data Drw using a single compression method, it is possible to achieve a higher compression ratio for the real-world data Drw.

[0067] (Decompression Process) Next, the decompression process performed by the information processing device 100 of this embodiment will be described. Figure 9 is a flowchart of the decompression process. The decompression process is the process of decompressing the archive Dar generated by the compression process described above to obtain the real-world data DRW before compression. The decompression process is started when the user operates the operation input unit 140 of the information processing device 100 and inputs a start command.

[0068] First, the data acquisition unit 112 (Figure 2) of the information processing device 100 acquires the archived Dar to be processed (S310). The data acquisition unit 112 may acquire the archived Dar from an external source via the interface unit 150. The data acquisition unit 112 may also acquire the archived Dar by generating it itself. The acquired archived Dar is stored, for example, in the storage unit 120.

[0069] Next, the decompression execution unit 118 (Figure 2) of the information processing device 100 decompresses the archived Dar to obtain the compressed partial data, the remaining partial data, and metadata Prw (S320). The decompression execution unit 118 refers to the metadata Prw and performs the decompression process on each partial data (S330). The decompression execution unit 118 merges the decompressed partial data to generate a single real-world data Drw (S340).

[0070] (Performance Evaluation) The performance of the real-world data compression and decompression method in the above-described embodiment was evaluated. Evaluations were conducted using real-world data generated by simulation, and evaluations were also conducted using actual real-world data.

[0071] The following types and order of data were used as real-world data for the simulation: • Type: Integer (INT), Floating-point number (FLOAT), String (STRING), Date and time (DATETIME) • Order: (Partially) Sequential, Random, Enumeration-Sequential, Enumeration-Random

[0072] For the actual real-world data used, we utilized publicly available online real-world data and real-world data owned by the Kawaguchi Laboratory at Nagoya University. This real-world data includes a variety of formats, such as CSV and JSON, sizes ranging from several hundred KB to several hundred MB, and data structures ranging from simple lists to more complex hierarchical structures. Table 2 shows an overview of the real-world data used.

[0073]

[0074] This performance evaluation investigated whether semantic metadata for real-world data could be generated using the LLM model. Therefore, experiments were conducted using the model shown in Table 3. The price in Table 3 is US dollars per million tokens. The target semantic metadata consists of the following four elements shown in Table 1: • dbp:itemType • schema:unitText • dbp:VariableCharacteracteristicEnumeration • dbp:dbpaDateTimeFormat

[0075]

[0076] The compression and decompression methods of the embodiments described above were used as examples. The methods of the examples are labeled "dbpa" and "dbpla" in each figure. As general data compression methods for comparative examples, bzip2, bzip3, gzip (Deflate), xz (LZMA2), and zstd were used. These methods are labeled "bz2", "bz3", "gz", "xz", and "zst" in each figure, respectively. As serialization methods for comparative examples, Amazon Ion (Binary) and JSON Bin Pack were used. These methods are labeled "ion" and "jbp" in each figure, respectively. However, since the compressed data size does not decrease significantly by simply applying the above serialization methods, experiments were conducted by combining these methods with xz compression.

[0077] For this performance evaluation, a desktop PC equipped with an Intel Core i9-12900K CPU and two DDR5-5600 48GB memory modules (total 96GB) was used. The operating system installed on this desktop PC was Ubuntu 22.04.3. The archiver versions used in this performance evaluation are as follows: bzip2 version 1.0.8, bzip3 version 1.5.1.r2, gzip version 1.10, xz version 5.2.5, zstd version 1.4.8, ion version 0.9.1, and jsonbinpack version 1.1.5.

[0078] Further details of the performance evaluation are as follows: • The evaluation of metadata generation for real-world data was performed using only actual real-world data. • For lossy compression, parameters were set so that the numerical error before and after compression was within 0.001%. • Each operation was repeated 64 times, and the mean and variance of the results were recorded. • In the comparative example, the options for specifying the compression ratio and the number of parallel processes were set to the highest compression ratio and the highest number of parallel processes.

[0079] The evaluation criteria are as follows: (1) Evaluation Item 1: Can the structural metadata of real-world data be correctly generated? (2) Evaluation Item 2: Can the semantic metadata of real-world data be accurately generated? (3) Evaluation Item 3: When real-world data is compressed based on metadata as in the example, can a higher compression ratio be achieved compared to conventional lossless compression methods? (4) Evaluation Item 4: When real-world data is compressed based on metadata as in the example, how does the compression / decompression speed compare to conventional lossless compression methods?

[0080] (Regarding evaluation item (1)) Regarding evaluation item (1) above, it was demonstrated that the method of the example can correctly generate structural metadata for all real-world data.

[0081] (Regarding evaluation item (2)) For evaluation item (2) above, the accuracy of semantic metadata generation was evaluated by comparing the metadata generated for each property in the @graph section with metadata provided by humans. The accuracy of each model and metadata property was calculated by dividing the number of instances in which the correct metadata matched the model output by the total number of occurrences of metadata across all real-world data. Furthermore, to evaluate the variability between different real-world data, the accuracy at each real-world data level was calculated and the standard error was determined.

[0082] Figure 10 is an explanatory diagram showing the evaluation results regarding the accuracy of semantic metadata generation. Figure 10 shows the accuracy when metadata was generated using each model shown in Table 3 for the four elements of semantic metadata described above. As shown in Figure 10, an accuracy of approximately 80% was obtained for all properties.

[0083] (Regarding evaluation item (3)) (Regarding evaluation using simulation data) In the evaluation using simulation data, a total of 30 different data patterns were used. These data are classified into the four data types and 26 subcategories mentioned above based on the data type and arrangement. For each type, the result displayed as "_ALL" in the figure represents the aggregated result including all subcategories and corresponds to the average performance of each type.

[0084] The following nine data arrangements were designed to mimic patterns commonly found in real-world data. The titles of the graphs in each figure follow the naming convention type_arrangement, indicating the data type and subcategory of arrangement. ・_CONST: A 4-column dataset where all values ​​are constant. ・_SEQ0: A 4-column dataset with sequential values. Each column has a 0%, 1%, 25%, and 90% probability of duplicate consecutive values ​​(e.g., 1, 2, 3, 3, 4, 5). STRING data is excluded. ・_SEQ2: A 4-column dataset with sequential values. 2% of the sequential values ​​are skipped (e.g., 1, 2, 3, 5, 6, 7). Duplicates occur with the same probability as in _SEQ0. STRING data is excluded. ・_SEQ10: A 4-column dataset with consecutive values. 10% of all values ​​are skipped. Duplicates occur with the same probability as in _SEQ0. STRING data is excluded. _RANDOM: A 4-column dataset where all values ​​are random. _enum64-random: A 4-column dataset where all values ​​are random, but the number of unique values ​​is limited to 64. DATETIME data is excluded. _enum64-seq0, _enum64-seq2, _enum64-seq10: Datasets with the same sequences as _SEQ0, _SEQ2, and _SEQ10. STRING and DATETIME data are excluded.

[0085] Figures 11 to 13 are explanatory diagrams showing the evaluation results of data compression ratio using simulation data. The compression methods in the examples (dbpa, dbpla) achieved the smallest archive size for most patterns. For INT and FLOAT type data, the numeric encoding layer effectively reduced the archive size. Sequential values ​​of INT, FLOAT, and DATETIME types benefited from exception-based encoding, resulting in further improved compression ratios. Furthermore, for the ENUM64 pattern, consistently good results were achieved for all data types.

[0086] (Regarding evaluation using actual data) Figures 14 to 16 are explanatory diagrams showing the evaluation results of data compression ratio using actual data. The compression methods of the examples (dbpa, dbpla) achieved the minimum size for 21 out of 27 datasets and outperformed the compression method of the comparative example for 23 datasets.

[0087] Figure 17 is an explanatory diagram showing the evaluation results of the average compression ratio when using actual data. For the average compression ratio of CSV files, the compression methods of the examples (dbpa, dbpla) showed a 13.6% improvement in compression ratio compared to bzip3. Among the examples, using DBPLA showed an improvement of 22.7%. For the average compression ratio of JSON files, the compression methods of the examples showed a 26.3% improvement in compression ratio compared to bzip3. Thus, it was confirmed that the compression methods of the examples achieve higher compression ratios than conventional methods for a large amount of real-world data.

[0088] (Regarding evaluation item (4)) (Regarding compression speed) Figure 18 is an explanatory diagram showing the evaluation results regarding compression speed for simulation data. As shown in Figure 18, the compression methods of the embodiment (dbpa, dbpla) achieved the fastest compression speed for many data types of real-world data. This result is thought to be due to the following factors. First, dictionary-based compression methods require additional time to generate the dictionary during compression. In contrast, the compression methods of the embodiment utilize metadata and the structure of real-world data, significantly reducing compression time. Second, integer, floating-point, and date / time data are represented as strings in CSV files, and are therefore usually processed as strings in conventional methods. However, the compression methods of the embodiment convert these values ​​to numerical format before compression, reducing the amount of data to be processed and improving processing speed. Finally, the ability to decompose real-world data into structural components and process them in parallel is also thought to have contributed to the speed improvement.

[0089] Figure 19 is an explanatory diagram showing the evaluation results for average compression speed using actual data. When using real-world data, the compression methods of the examples (dbpa, dbpla) were significantly faster than xz and zstd. On the other hand, the compression methods of the examples were about twice as slow as bzip2, bzip3, and gzip. This result suggests that when the original data size is small, overhead occurs in the structure decomposition using metadata and the application of compression layers, which may negate the advantages of the methods of the examples. However, in practical applications involving large-scale data, this overhead can be ignored.

[0090] (Regarding decompression speed) Figure 20 is an explanatory diagram showing the evaluation results for decompression speed using simulation data. Figure 21 is an explanatory diagram showing the evaluation results for average decompression speed using actual data. As shown in Figures 20 and 21, the decompression speed when using the decompression methods of the embodiment (dbpa, dbpla) was equivalent to the decompression speed when using existing methods. In particular, the decompression methods of the embodiment achieved a speed approximately three times faster than bzip3, which has a high compression ratio. It is thought that the ease with which the methods of the embodiment can be processed in parallel contributed greatly to the improvement in decompression speed.

[0091] (Effects of this embodiment) As described above, the information processing device 100 of this embodiment comprises a data acquisition unit 112, a data splitting unit 113, a method selection unit 114, and a compression execution unit 115. The data acquisition unit 112 acquires real-world data DRW. The data splitting unit 113 splits the real-world data DRW into a plurality of partial data. The method selection unit 114 assigns a compression method to each partial data based on the characteristics of each partial data. The compression execution unit 115 compresses each partial data using the assigned compression method. The aggregation unit 116 combines the compressed partial data to generate a single archive Dar.

[0092] In this embodiment, for each of the multiple partial data contained in the real-world data DRW, a compression process is performed using a compression method assigned based on the characteristics of the partial data, and the compressed partial data is combined into a single archive DAR. Therefore, compared to an embodiment in which a single compression method is used to compress the entire real-world data, a higher compression ratio can be achieved for the real-world data DRW. For example, even among structured data, there is a method that uses different types of compression for each column in a column-oriented database for tabular data, but storing data in a column-oriented database itself is more expensive than storing data in normal storage. In this embodiment, a high compression ratio can be achieved while avoiding such an increase in data storage costs.

[0093] In this embodiment, the characteristics of the partial data may include at least one of the data type and the data arrangement. In this way, an appropriate compression method can be assigned to each partial data based on at least one of the data type and arrangement, thereby achieving compression of real-world data DRW with a higher compression ratio.

[0094] In this embodiment, the data acquisition unit 112 may acquire metadata Prw that identifies the characteristics of each subdata contained in the real-world data Drw, and the method selection unit 114 may assign a compression method based on the metadata Prw. In this way, the characteristics of the subdata can be grasped quickly and reliably, and an appropriate compression method can be assigned to each subdata, thereby achieving compression of the real-world data Drw with a higher compression ratio.

[0095] In this embodiment, the aggregation unit 116 may generate an archive Dar that includes metadata Prw. In this way, it is possible to generate an archive Dar that includes each portion of data compressed with a high compression ratio and also includes metadata Prw, thereby obtaining an archive Dar that can be used for data utilization such as data conversion, analysis, and visualization.

[0096] In this embodiment, the data acquisition unit 112 may generate metadata Prw based on real-world data Drw. In this way, even real-world data Drw that does not yet have metadata Prw attached can be compressed with a high compression ratio.

[0097] In this embodiment, the data acquisition unit 112 may analyze the real-world data DRW to create structural metadata Pst, create semantic metadata Pse based on the structural metadata Pst, and merge the structural metadata Pst and semantic metadata Pse to generate metadata PRW. In this way, metadata PRW can be attached to the real-world data DRW with high accuracy.

[0098] In this embodiment, the information processing device 100 further includes a storage unit 120 that stores correspondence relationship information CR, which indicates the correspondence between data characteristics and compression methods, and the method selection unit 114 may assign a compression method by referring to the correspondence relationship information CR. In this way, an appropriate compression method can be assigned to each portion of data quickly and reliably, and compression of real-world data DRW with a high compression ratio can be achieved.

[0099] In this embodiment, the information processing device 100 may further include an information update unit 117 that updates the correspondence information CR. In this way, as new compression methods are developed, a more appropriate compression method can be assigned to each part of the data, and a higher compression ratio can be achieved for the real-world data DRW.

[0100] In this embodiment, the correspondence information CR may be information that shows the correspondence between data characteristics, compression methods, and index values ​​that correlate with the compression ratio. In this way, an appropriate compression method can be quickly and reliably assigned to each portion of data according to the required data accuracy, and compression of real-world data DRW with a high compression ratio can be achieved.

[0101] In this embodiment, the information processing device 100 may further include an extraction execution unit 118 that extracts the archived Dar to obtain the real-world data DRW before compression. In this way, it is possible to extract the real-world data DRW that has been compressed with a high compression ratio.

[0102] (Modifications) The technologies disclosed herein are not limited to the embodiments described above and can be modified in various forms without departing from the spirit thereof, for example, the following modifications are possible.

[0103] The configuration of the information processing device 100 in the above embodiment is merely an example and can be modified in various ways. Furthermore, the content of each process in the above embodiment is merely an example and can be modified in various ways. For example, the configuration of the metadata Prw in the above embodiment can be modified in various ways.

[0104] In the above embodiment, a compression method is assigned to each combination of data type and arrangement in the correspondence relationship information CR, but a compression method may also be assigned to other characteristics of the data.

[0105] In the above embodiment, at least one of the functional units included in the control unit 110 of the information processing device 100 may be included in a device other than the control unit 110 of the information processing device 100. Each process in the above embodiment does not necessarily have to be executed by a single device, but may be executed by different devices. Each step of each process in the above embodiment does not necessarily have to be executed by a single device, but may be executed by different devices. In the above embodiment, a part of the configuration implemented by hardware may be replaced with software, and conversely, a part of the configuration implemented by software may be replaced with hardware.

[0106] The technologies disclosed herein are not limited to real-world data but are equally applicable to structured data in general. Examples of structured data other than real-world data include: • API responses published on the following website: https: / / apidog.com / jp / blog / how-to-use-json-in-api-response / • JSON-LD based on schema.org published on the following website: https: / / keywordmap.jp / academy / seo-structured-data /

[0107] 100: Information processing unit 110: Control unit 111: Data processing unit 112: Data acquisition unit 113: Data splitting unit 114: Method selection unit 115: Compression execution unit 116: Aggregation unit 117: Information update unit 118: Decompression execution unit 120: Storage unit 130: Display unit 140: Operation input unit 150: Interface unit 190: Bus CP: Data processing program CR: Correspondence information Drw: Real-world data Prw: Metadata Pse: Semantic metadata Pst: Structural metadata

Claims

An information processing device, A data acquisition unit that acquires structured data, A data splitting unit that divides the structured data into multiple subdata, A method selection unit assigns a compression method to each of the aforementioned partial data based on the characteristics of each of the aforementioned partial data, A compression execution unit that compresses each of the aforementioned partial data using an assigned compression method, An aggregation unit that combines each of the compressed partial data to generate a single archive, An information processing device equipped with the following features.   An information processing apparatus according to claim 1, The aforementioned characteristics include at least one of the data type and the data arrangement.   An information processing apparatus according to claim 1 or claim 2, The data acquisition unit acquires metadata that identifies the characteristics of each of the subdata included in the structured data, The method selection unit is an information processing device that assigns the compression method based on the metadata.   An information processing apparatus according to claim 3, The aggregation unit is an information processing device that generates the archive including the metadata.   An information processing apparatus according to claim 3 or claim 4, The data acquisition unit is an information processing device that generates metadata based on the structured data.   An information processing device according to claim 5, The data acquisition unit, The structured data is analyzed to create structured metadata from the metadata of the structured data. Based on the aforementioned structured metadata, semantic metadata is created from the metadata of the structured data. An information processing device that generates metadata by merging the structural metadata and the semantic metadata.   An information processing device according to any one of claims 1 to 6, The system further includes a storage unit that stores correspondence information indicating the correspondence between the aforementioned characteristics and the aforementioned compression method. The method selection unit is an information processing device that assigns the compression method by referring to the correspondence information.   An information processing apparatus according to claim 7, An information processing device further comprising an information update unit that updates the aforementioned correspondence information.   An information processing apparatus according to claim 7 or claim 8, The aforementioned correspondence information is information that shows the correspondence between the characteristics, the compression method, and an index value that correlates with the compression ratio, in an information processing device.   An information processing apparatus according to any one of claims 1 to 9, An information processing device further comprising a decompression execution unit that decompresses the archive to obtain the structured data before compression.   Information processing method, Obtain structured data, The aforementioned structured data is divided into multiple subdata sets, Based on the characteristics of each of the aforementioned partial data, a compression method is assigned to each of the aforementioned partial data. Each of the aforementioned partial data is compressed using the assigned compression method. An information processing method for generating a single archive by combining the compressed partial data.   On the computer, The process of obtaining structured data, The process of dividing the aforementioned structured data into multiple subdata, A process of assigning a compression method to each of the aforementioned partial data based on the characteristics of each of the aforementioned partial data, A process to compress each of the aforementioned partial data using the assigned compression method, A process to combine each of the compressed partial data to generate a single archive, A computer program that executes something.