Information compression system and information compression method

The information compression system enhances compression efficiency for multi-dimensional sensor data by converting element values into identification information and compressing it, addressing the inefficiencies of existing techniques.

JP7689093B2Active Publication Date: 2025-06-05HITACHI LTD

Patent Information

Application Number
JP2022028562
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-25
Publication Date
2025-06-05
Estimated Expiration
2042-02-25

AI Technical Summary

Technical Problem

Existing data compression techniques for multi-dimensional sensor data, such as those used in urban space design and mobility, do not achieve sufficient compression efficiency.

Method used

An information compression system that acquires data, determines objects and their meanings, and generates compression target data by converting element values into identification information, which is then compressed to produce higher compression efficiency.

Benefits of technology

The system achieves higher compression efficiency by converting high-randomness pixel values into low-randomness identification information, allowing for a higher compression rate while maintaining data integrity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007689093000001
    Figure 0007689093000001
  • Figure 0007689093000002
    Figure 0007689093000002
  • Figure 0007689093000003
    Figure 0007689093000003
Patent Text Reader

Abstract

To provide an information compression system capable of achieving higher compression efficiency.SOLUTION: A data acquisition section 101 acquires data. A generation section (segmentation section 102 and integration section 103) determines each object depicted by the data and a sense of each object and generates compression target data obtained by converting values of elements in the data to identification information representing each object and the sense of each object according to a determination result. A data storage section 104 generates compressed data obtained by compressing the compression target data. That makes it possible to convert values of highly random elements to slightly random identification information to compress the converted information while reducing the amount of information. Therefore, the compression ratio can be increased.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information compression system and an information compression method.

Background Art

[0002] In recent years, in fields such as urban space design, social infrastructure, and mobility, sensor data such as multi-dimensional point cloud data and image data has been acquired by sensors such as LiDAR (Light Detection and Ranging) and cameras, and is utilized for various applications. However, in these fields, there is a problem that the amount of sensor data becomes enormous.

[0003] On the other hand, Patent Document 1 discloses a technique for compressing multi-dimensional data using a neural network. According to this technique, optimal compression is possible regardless of the number of dimensions and format of the multi-dimensional data.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the technique described in Patent Document 1, since the multi-dimensional data is compressed as it is, sufficient compression efficiency may not be obtained.

[0006] An object of the present disclosure is to provide an information compression system and an information compression method capable of achieving higher compression efficiency.

Means for Solving the Problems

[0007] An information compression system according to an aspect of the present disclosure is an information compression system that compresses data, and includes an acquisition unit that acquires the data, a generation unit that determines each object depicted in the data and the meaning of each object, and based on the determination result, generates compression target data in which the values of the elements of the data are converted into identification information representing each object and the meaning of each object, and a compression unit that generates compressed data by compressing the compression target data.

Effect of the Invention

[0008] According to the present invention, it becomes possible to achieve higher compression efficiency.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Mode for Carrying Out the Invention

[0010] Hereinafter, examples of the present disclosure will be described with reference to the drawings.

Example

[0011] FIG. 1 is a diagram showing the physical configuration of the information compression system according to Example 1 of the present disclosure. The information compression system shown in FIG. 1 is configured by Node 1. Node 1 is communicably connected to Sensor 2 and Input / Output Device 3. Also, there may be a plurality of Nodes 1. When there are a plurality of Nodes 1, at least one of the Nodes 1 may be connected to Sensor 2 and Input / Output Device 3. In the example of FIG. 1, two Nodes 1 are communicably connected to each other via Network 4, and Sensor 2 is connected to one of the Nodes 1.

[0012] Node 1 is a computer system, for example, a cloud system, an on-premises system, edge computing, or a mobile device such as a smartphone.

[0013] Node 1 has a main processor 11, a main memory 12, a storage 13, IFs 14 and 15, and a subprocessing unit 16, and they are interconnected via an internal bus 17.

[0014] The main processor 11 is, for example, a CPU (Central Processing Unit), etc., and executes various processes according to a program by reading the program (computer program) from the storage 13 into the main memory 12. The main memory 12 is a storage device used as a work area for the program. The storage 13 is a storage device that stores programs and information used and generated by the main processor 11 and the subprocessing unit 16. IFs 14 and 15 are communication devices that are communicably connected to external devices. In FIG. 1, IF 14 is connected to the sensor 2, and IF 15 is connected to another node 1 via the network 4.

[0015] The subprocessing unit 16 is a unit for executing predetermined processes according to a program, and is, for example, a GPU (Graphics Processing Unit), etc. The subprocessing unit 16 has a plurality of cores 16a for performing multiprocessing that executes a plurality of processes simultaneously, and a submemory 16b used as a work area by the cores 16a.

[0016] The sensor 2 is various sensing devices such as an optical sensor like Lidar, an optical camera, and a gravity sensor, and transmits the detected sensor data to the node 1. Note that instead of or in addition to the sensor 2, a recording medium storing sensor data such as an SD card may be used.

[0017] The input / output device 3 includes an input device that receives various information from a user who uses an information compression system such as a keyboard, a touch panel, and a pointing device, and an output device that outputs various information to the user such as a display device and a printer. Further, the input / output device 3 may be a mobile terminal used by the user.

[0018] FIG. 2 is a diagram showing the logical configuration of Node 1. As shown in FIG. 2, Node 1 includes a data acquisition unit 101, a segmentation unit 102, an integration unit 103, a data storage unit 104, a data conversion unit 105, and a data utilization unit 106. Each of the units 101 to 106 of Node 1 is realized, for example, by at least one of the main processor 11 and the sub-processing unit 16 executing a program.

[0019] The data acquisition unit 101 acquires sensor data from the sensor 2. In this embodiment, the sensor data includes three-dimensional point cloud data acquired by a surveying sensor such as a Lidar or TOF (Time of Flight) sensor, and color camera data which is image data acquired by a color camera. The point cloud data is a set of point data indicating each position on the surface of an object, and each point data includes coordinate information indicating the coordinates of the position on the surface of the object. The color camera data has color information indicating a plurality of colors (in this embodiment, R (red), G (green), and B (blue)) for each pixel as a pixel value. Also, the point cloud data and the color camera data are time-series data in this embodiment. Specifically, the color camera data is moving image data.

[0020] Note that the data acquisition unit 101 may provide a user with a data utilization interface for specifying the sensor data to be acquired, and acquire the sensor data specified by the data utilization interface. Instead of the data utilization interface, a command or an API (Application Programming Interface) or the like for specifying the sensor data to be acquired may be used.

[0021] The segmentation unit 102 determines each object shown in the color camera data acquired by the data acquisition unit 101 and the meaning of each object, and generates segmentation data in which each pixel value of the color camera data is converted into identification information representing each object and the meaning of each object based on the determination result. The identification information includes an instance ID which is identification information for identifying an object and a meaning ID which is identification information for identifying the meaning of the object. Note that in the segmentation data, at least a part of each pixel value may be converted into identification information.

[0022] For the segmentation process which is the process by the segmentation unit 102, for example, a segmentation model which is a learned model for determining each object shown in the color camera data and the meaning of each object, and a management table for defining the instance ID and the meaning ID are used. The segmentation unit 102 may provide a user with a setting interface for performing settings related to the segmentation process, and execute the segmentation process according to the settings via the setting interface.

[0023] The integration unit 103 generates integrated data by integrating the point cloud data acquired by the data acquisition unit 101 and the segmentation data which is the processing result by the segmentation unit 102 based on the sensor information regarding the sensor 2. Specifically, the integrated data has coordinate information indicating the coordinates of a position and identification information representing the object at that position for each position on the surface of the object shown in the color camera data. In this embodiment, the integrated data is data to be actually compressed, and the segmentation unit 102 and the integration unit 103 constitute a generation unit for generating the data to be compressed. Note that the integration unit 103 uses the sensor information 50 to convert the point cloud data acquired by the data acquisition unit 101 into a unified coordinate space (the coordinate space specified by the data utilization interface 80). Thereby, data can be uniformly handled by the data utilization interface 80, and when the same object is photographed by sensors with multiple viewpoints, redundant point cloud data can be efficiently reduced in data amount by quantization and compression processing described later.

[0024] The data storage unit 104 has the function of a compression unit that generates compressed data by compressing the integrated data integrated by the integration unit 103, and the function of an expansion unit that generates expanded data by expanding the compressed data.

[0025] For example, the data storage unit 104 compresses the integrated data in units of data blocks called chunks. Further, the data storage unit 104 may store the compressed data in its own node, which is node 1 having the data storage unit 104 that performed the compression, or may transfer the compressed data to another node, which is another node 1 different from its own node, and store it in the other node. Further, the data storage unit 104 reads and expands the compressed data in response to a predetermined timing or an instruction from the user.

[0026] The data conversion unit 105 converts the expanded data generated by the data storage unit 104 into data in a predetermined format. For example, the data conversion unit 105 converts the expanded data into mesh data or learning data for machine learning. Further, the expanded data may be used as it is without being converted.

[0027] The data utilization unit 106 provides the converted data converted by the data conversion unit 105 as output data. For example, the data utilization unit 106 displays the output data on the input / output device 3 or transmits it to another node 1 to achieve real-time utilization or time-series utilization of the expanded data. Real-time utilization is, for example, inference and visualization, and time-series utilization is learning and analysis.

[0028] Figure 3 is a diagram showing an example of sensor data. The sensor data 30 shown in Figure 3 has point cloud data 30a and color camera data 30b.

[0029] The point cloud data 30a is a set of point data 31 that is coordinate information indicating the positions on the surface of an object, and each point data indicates the coordinates of the position on the surface of the object in an orthogonal coordinate system defined by the x-axis, y-axis, and z-axis. The x-axis, y-axis, and z-axis may be defined for each sensor 2 or may be defined widely.

[0030] The color camera data 30b indicates the pixel value for each pixel arranged in a matrix in the two-dimensional directions of the horizontal axis and the vertical axis. The pixel value has color information consisting of a plurality of values representing different colors (in this embodiment, R (red), G (green), B (blue)). Therefore, the color camera data 30b can also be regarded as three-dimensional array data ([3][X][Y]). Here, [3] indicates color information, [X] indicates the pixel position in the vertical axis direction, and [Y] indicates the pixel position in the vertical axis direction. In FIG. 3, for simplicity, only the pixel values corresponding to a single color are shown.

[0031] In this embodiment, the point cloud data 30a and the color camera data 30b are time-series data, and in FIG. 3, the point cloud data 30a and the color camera data 30b at a certain point in time are shown.

[0032] FIG. 4 is a diagram showing an example of integrated data. The integrated data 40 shown in FIG. 4 has fields 410 to 43. Field 41 stores coordinate information (point data) indicating the position of the surface of the object. Field 42 stores a meaning ID for identifying the meaning of the object at the position indicated by the coordinate information in Field 41. Field 43 stores an instance ID for identifying the object at the position indicated by the coordinate information in Field 41.

[0033] FIG. 5 is a diagram showing an example of sensor information. The sensor information 50 shown in FIG. 5 has fields 51 to 59.

[0034] Field 51 stores a sensor ID which is identification information for identifying sensor 2. Field 52 stores the type of sensor 2 identified by the sensor ID in Field 51. In this embodiment, the types include "point cloud" which is sensor 2 (for example, Lidar) for acquiring point cloud data, and "color camera" corresponding to sensor 2 (for example, color camera) for acquiring color camera data. Field 53 stores a pair ID for specifying a pair of sensor 2 for acquiring point cloud data and color camera data for generating integrated data. In the example of FIG. 5, in Field 53 corresponding to Field 52 storing "color camera" as the type, the sensor ID of sensor 2 of "point cloud data" which is paired with the "color camera" sensor 2 is stored as the pair ID.

[0035] Field 54 stores position information indicating the position where sensor 2 is arranged. The position information indicates the position of sensor 2 in a rectangular coordinate system of the x-axis, y-axis, and z-axis. Note that the coordinate axes (x-axis, y-axis, and z-axis) for defining the position of sensor 2 do not have to be the same as the coordinate axes of the point cloud data shown in FIG. 2. Field 55 stores orientation information indicating the orientation of sensor 2. In the example of FIG. 5, the orientation information is indicated by the rotation angle Ψ, elevation angle θ, and azimuth angle Φ. Field 56 stores the scale of sensor 2. Field 57 stores the focal length of sensor 2. Field 58 stores the resolution of sensor 2. Field 59 stores the field of view angle of sensor 2.

[0036] FIG. 6 is a diagram showing an example of chunk information regarding a chunk which is a data block unit for compressing integrated data. The chunk information 60 shown in FIG. 6 has fields 61 to 68.

[0037] Field 61 stores a sensor ID that identifies the sensor 2 used to generate the integrated data to be compressed. Field 62 stores the starting position in the x direction of a chunk in the integrated data, Field 63 stores the starting position in the y direction of a chunk in the integrated data, and Field 64 stores the starting position in the z direction of a chunk in the integrated data. Field 65 stores the start time of a chunk in the integrated data, and Field 66 stores the end time of a chunk in the integrated data. Note that the width of each chunk in the x direction, y direction, and z direction is specified in advance separately from, for example, the chunk information 60. Note that Field 61 stores a plurality of sensor IDs, and information acquired from a plurality of sensors may be stored in the same chunk.

[0038] Field 67 stores the compression state of a chunk. The compression state indicates whether the chunk has been compressed, and if the chunk has been compressed, it further indicates the compression algorithm used to compress the chunk. Field 68 stores the compressed data obtained by compressing the chunk. The compressed data includes, for example, the compressed binary data that is the compressed chunk body, reference information indicating the management table used for compressing the chunk, and setting values related to the normalization performed during compression. The reference information is, for example, a pointer indicating the management table. The setting values are, for example, the minimum value and the maximum value corresponding to each coordinate axis when normalization is performed using the Min-Max method.

[0039] Note that in the example of FIG. 6, the chunks are set by position and time, but the method of setting chunks is not limited to this example. For example, the chunks may be set according to at least one of the instance ID and the meaning ID.

[0040] FIG. 7 is a diagram showing an example of a management table. The management table 70 shown in FIG. 7 has a meaning management table 70a and an instance management table 70b.

[0041] The meaning management table 70a has fields 71 and 72. Field 71 stores a meaning ID. Field 72 stores meaning information indicating the meaning identified by the meaning ID. In this embodiment, the meaning information indicates, as a meaning, the type of an object such as "person" or "desk".

[0042] The instance management table 70b has fields 73 and 74. Field 73 stores an instance ID. Field 74 stores a wide-area ID that widely identifies the object identified by the instance ID. Note that the instance ID is identification information for identifying an object in a single integrated data (or a single target space), and the wide-area ID is identification information for identifying an object common to all integrated data.

[0043] FIG. 8 is a diagram showing an example of a data utilization interface for reading output data via the data utilization unit 106. The data utilization interface 80 shown in FIG. 8 includes designated columns 81 to 87.

[0044] Designated column 81 is a column for designating a sensor ID that identifies sensor 2 corresponding to the output data to be read. Designated column 82 is a column for designating the spatial start position of the sensor data to be acquired, and designated column 83 is a column for designating the spatial end position of the output data to be read. In the example of FIG. 8, the start position and the end position designate x, y, z coordinates. Designated column 84 is a column for designating the start time of the output data to be read, and designated column 85 is a column for designating the end time of the output data to be read. Designated column 86 is a column for designating a meaning ID representing the meaning of the output data to be read, and designated column 87 is a column for designating the instance of the output data to be read. When the meaning ID and the instance ID are designated, only the output data corresponding to the designated ID is acquired.

[0045] In the specified columns 81 to 87, it is also possible to set "Any" which specifies all. Also, "Real Time" can be specified for the start time or the end time. In this case, the output data corresponding to the sensor data of the current time acquired by the data acquisition unit 101 is read out in real time as a stream via the data utilization unit 106.

[0046] FIG. 9 is a diagram showing an example of a setting interface for performing settings related to the segmentation process by the segmentation unit 102. The setting interface 90 shown in FIG. 9 has selection columns 91 to 93, and setting buttons 94 and 95.

[0047] The selection column 91 is a column for specifying the storage location of the segmentation model used in the segmentation process. The selection column 92 is a column for specifying the storage location of the management table used in the segmentation process. The selection column 93 is a column for specifying the conversion content and the necessity of acquisition for converting color camera data into segmentation data in the segmentation process. Specifically, in the selection column 93, the meaning of the object for which data is to be acquired and whether to convert the pixel value of the pixel into identification information or leave it as color information are specified. For example, if it is set that acquisition is not required for the meaning "desk", the data portion determined to be "desk" in the result of segmentation is deleted and not stored. By this function, acquisition of unnecessary data can be suppressed and the storage capacity can be saved.

[0048] The setting button 94 is a button for setting the segmentation model and the management table. When pressed, the segmentation model and the management table stored in the storage locations specified in the selection columns 91 and 92 are set. The setting button 95 is a button for setting the conversion content. When pressed, the conversion content is set.

[0049] By using the setting interface, it becomes possible to delete color information for an object having a specific meaning such as "person" and replace it with a meaning ID and an instance ID, so that privacy can be protected.

[0050] FIG. 10 is a flowchart for explaining an example of a writing process that is a process until sensor data is compressed and stored.

[0051] In the writing process, first, the data acquisition unit 101 acquires sensor data from the sensor 2 (step S101). Here, the sensor data includes point cloud data and color camera data.

[0052] The segmentation unit 102 analyzes the color camera data in the sensor data acquired by the data acquisition unit 101 using the segmentation model and the management table set in the setting interface, and identifies the object shown in the color camera data and its meaning. Then, the segmentation unit 102 acquires an instance ID that is identification information for identifying the identified object and a meaning ID that is identification information for identifying the meaning of the object (step S102).

[0053] The segmentation unit 102 converts the color camera data into segmentation data according to the conversion content set in the setting interface 90 (step S103). As a result, filtering is performed so that only the pixels showing the object having the meaning set in the setting interface 90 remain in the segmentation data as identification information or color information.

[0054] The integration unit 103 generates integrated data by integrating corresponding point cloud data and segmentation data based on the sensor information (step S104). Here, when acquiring point cloud data and color camera data as sensor data, the integration unit 103 generates integrated data with identification information attached to the coordinate points of the point cloud data corresponding to the spatial positions of the pixels of the segmentation data. Also, even when the data acquisition unit 101 does not acquire point cloud data, there is a method for generating integrated data. Specifically, the integration unit 103 obtains segmentation data from the color camera data, calculates a depth map (the distance from the sensor of the object corresponding to the pixel) by a general method such as a depth estimation method, and calculates spatial coordinate points from the calculated depth map. Thereby, the integration unit 103 can obtain information similar to the point cloud data and use it to generate integrated data having coordinate point information in the same manner.

[0055] The data storage unit 104 divides the integrated data generated by the integration unit 103 into a plurality of chunk data and generates chunk information regarding each chunk data (step S105).

[0056] The data storage unit 104 normalizes each chunk data (step S106). Here, the data storage unit 104 normalizes each chunk data using the Min - Max method.

[0057] The data storage unit 104 determines whether to perform synchronous compression for compressing the chunk data at this timing (step S107). Note that whether to perform synchronous compression is, for example, preset.

[0058] When performing synchronous compression (step S107: Yes), the data storage unit 104 quantizes each chunk of data (step S108). Here, quantization means, for example, for point cloud data represented by floating-point coordinates, dividing by a value called the quantization width and then integerizing it by an operation such as the round function. When normalizing, the quantization width can be used to adjust the granularity of quantization. For example, in the case of a point cloud, after quantization, there may be duplicate identical coordinate points, and by deleting them, the data volume can be efficiently reduced. Here, the deletion of duplicate identical coordinate points can be executed at high speed on the subprocessing unit 16, for example, by using a unique function or the like that eliminates duplicate elements in the processing system of machine learning. Also, when the granularity of quantization is coarsened, the accuracy decreases and the data volume decreases, and when the granularity is made finer, the accuracy increases and the data volume increases. That is, by adjusting the granularity (quantization width) of quantization, the balance between the data volume and the accuracy of the coordinates can be adjusted. Furthermore, the quantization and deletion of identical coordinate points described here can efficiently eliminate more duplicate elements by simultaneously processing sensor data obtained by photographing the same object from multiple viewpoints, so that the overall data volume can be efficiently reduced. Then, the data storage unit 104 generates compressed data obtained by compressing each quantized chunk of data as target data (step S109). On the other hand, when synchronous compression is not performed (step S107: No), the data storage unit 104 skips the processes of steps S108 and S109 and uses the chunk data as the target data.

[0059] Then, the data storage unit 104 determines whether to perform synchronous transfer to transfer each target data to other nodes at this timing (step S110). Note that whether to perform synchronous transfer is, for example, preset.

[0060] When performing synchronous transfer (step S110: Yes), the data storage unit 104 transfers the target data to another node 1 (step S111). Then, when the data storage unit 104 of another node 1 receives the target data, it stores the target data (step S112) and ends the writing process. Also, when not performing synchronous compression (step S110: No), the data storage unit 104 skips the process of step S111, stores the target data in the own node which is this node 1 (step S112), and ends the writing process.

[0061] Each process (steps S101 to S112) of the writing process described above may be executed on separate nodes 1. In this case, a transfer process for transferring data to another node 1 is performed between each process. Also, each chunk data for which synchronous compression was not performed can be compressed at an arbitrary timing, and the target data for which synchronous transfer was not performed can be transferred to another node 1 at an arbitrary timing.

[0062] FIG. 11 is a flowchart for explaining an example of a reading process which is a process until decompressing and outputting compressed data.

[0063] In the reading process, the data storage unit 104 specifies the chunk data to be decompressed as the target chunk data (step S201). For example, the data utilization unit 106 provides a data utilization interface to allow the user to specify the chunk data to be decompressed, and the data storage unit 104 specifies the chunk data specified by the user as the target chunk data.

[0064] The data storage unit 104 determines whether the target chunk data is stored in the own node (step S202).

[0065] If the target chunk data is stored in its own node (step S202: Yes), the data store unit 104 reads the target chunk data (step S203). On the other hand, if the target chunk data is not stored in its own node (step S202: No), the data store unit 104 reads the target chunk data from another node 1 that stores the target chunk data (step S204).

[0066] Then, the data store unit 104 determines whether the read target chunk data is compressed (step S205).

[0067] If the target chunk data is compressed (step S205: Yes), the data store unit 104 decompresses the target chunk data (step S206). The data store unit 104 performs inverse quantization on the decompressed target chunk data (step S207) and further performs renormalization (step S208). The inverse quantization here means multiplying the quantization width at the time of compression to the target chunk data to return to the scale of the original value. If the target chunk data is not compressed (step S205: No), the data store unit 104 skips the processes of steps S206 to S208.

[0068] Then, the data store unit 104 combines the target chunk data to generate extended data (step S209). Note that when the compression of the chunk data is lossless compression, the extended data becomes integrated data.

[0069] The data conversion unit 105 converts the extended data generated by the data store unit 104 into data of a predetermined format and outputs it (step S210), and ends the read process.

[0070] Each process of the read process (steps S201 to S210) described above may be executed on separate nodes 1. In this case, a transfer process for transferring data to other nodes 1 is performed between each process.

[0071] FIG. 12 is a diagram showing a more detailed configuration of the data store unit 104. The data store unit 104 includes, as a configuration for compression processing, a normalization / quantization unit 201, a voxelizer 202, an entropy estimator 203, and an entropy encoder 204, and includes, as a configuration for decompression processing, an entropy decoder 211, an entropy estimator 212, a PC (Point Cloud) converter 213, and an inverse quantization / renormalization unit 214. The entropy estimators 203 and 212 may have the same configuration.

[0072] In the compression process, first, the normalization / quantization unit 201 performs normalization and quantization on the coordinate information in the chunk data. The normalization and quantization are performed for each of the coordinate axes (x-axis, y-axis, and z-axis) defining the three-dimensional space.

[0073] Subsequently, the voxelizer 202 generates voxel information obtained by voxelizing the quantized chunk data, which is the chunk data whose coordinate information has been normalized and quantized. Specifically, the voxelizer 202 divides the chunk data into a plurality of voxels, which are three-dimensional regions having a predetermined volume, and sets the value of each voxel based on the identification information (semantic ID and instance ID) corresponding to each coordinate included in the voxel. Specifically, the value of each voxel is set to the semantic ID and instance ID having the largest number among the semantic ID and instance ID corresponding to each coordinate included in the voxel. As a result, the chunk data is converted into voxel information, which is a set of a voxel Ch1 having a semantic ID (S) as a value and a voxel Ch2 having an instance ID (I) as a value.

[0074] Note that when color information is associated with each coordinate information instead of the semantic ID and instance ID, the voxel information becomes a set of a voxel Ch3 having a value indicating red (R), a voxel Ch4 having a value indicating green (G), and a voxel Ch5 having a value indicating blue (B). Also, each voxel may be represented by an octree structure.

[0075] The entropy estimator 203 estimates the entropy of the voxel information. Here, the entropy estimator 203 estimates, as entropy, a probability distribution (hereinafter, sometimes simply referred to as a probability distribution) representing the occurrence probability of each symbol that can be the value of the voxel information. The entropy estimator 203 is constructed, for example, with a learned model using a DNN (Deep Neural Network) such as a multi-layer 3D convolutional neural network (CNN). The entropy estimator 203 may take as input voxel information with low resolution and estimate the probability distribution of voxel information with high resolution. In this case, for improving the prediction accuracy and coping with decoding of various resolutions (generally called progressive, etc.), voxel information of a plurality of resolutions may be stepwise used as the input to the entropy estimator 203 to estimate the probability distribution. Also, for improving the estimation accuracy, when dealing with time-series data, past voxel information with high similarity or voxel information subjected to statistical processing (for example, processing for obtaining the median, average value, or variance of a predetermined period) may be used as the input. Further, for improving the estimation accuracy, by using a multi-layer 3D CNN as an autoregressive model and using the symbol values of known voxel information such as in the vicinity of the estimation target as the input to the entropy estimator 203, the probability distribution of the symbol of the estimation target may be estimated. Also, for the plurality of input data to the entropy estimator 203 corresponding to the plurality of methods described above, after matching the data resolutions, they may be combined, for example, by combining them by inputting them into the channels of a multi-layer 3D CNN to improve efficiency.

[0076] Then, the entropy encoder 204 encodes the voxel information based on the probability distribution estimated by the entropy estimator 203 to generate compressed binary data.

[0077] In the expansion process, the entropy decoder 211 decodes the compressed binary data to generate voxel information. Specifically, the entropy decoder 211 uses the entropy estimator 212 to predict the probability distribution of the voxel values (symbols), and the entropy decoder 211 decodes the symbols using the compressed binary data and the predicted probability distribution, and finally decodes them as voxel information. As the entropy estimator 212, the same one as the entropy estimator 203 is used to obtain the estimation result of the same probability distribution during encoding. Also, the input to the entropy estimator 212 is the same as the input to the entropy estimator 203 during encoding. Also, when performing step-by-step probability distribution estimation at different resolutions during compression or probability distribution estimation using an autoregressive model, the probability distribution estimation by the entropy estimator 212 and the voxel information decoding by the entropy decoder 211 are repeated multiple times to finally decode the voxel information.

[0078] The PC converter 213 converts the voxel information generated by the entropy decoder 211 into quantization chunk data having coordinate information and identification information. The inverse quantization / renormalizer 214 performs inverse quantization and renormalization on the coordinate information of the quantization chunk data converted by the PC converter 213 to generate chunk data as expansion data.

[0079] As described above, according to this embodiment, the data acquisition unit 101 acquires color camera data. The generation unit (segmentation unit 102 and integration unit 103) determines each object shown in the color camera data and the meaning of each object, and based on the determination result, generates compression target data obtained by converting the pixel values of the color camera data into identification information representing each object and the meaning of each object. The data storage unit 104 generates compressed data obtained by compressing the compression target data. Therefore, while reducing the amount of information, it is possible to convert high-randomness pixel values into low-randomness identification information and compress them, so that a higher compression rate can be achieved.

[0080] Also, in this embodiment, the data acquisition unit 101 further acquires point cloud data having a plurality of pieces of coordinate information indicating each position on the surface of the object, and the generation unit generates integrated data having coordinate information and identification information representing the object at the position for each position on the surface as data to be compressed. Therefore, it is possible to achieve a higher compression rate. Furthermore, the data acquisition unit 101 may acquire, as the data to be acquired, in addition to an image having two dimensions of vertical and horizontal, three-dimensional voxel data having vertical, horizontal, and depth dimensions. In that case, the segmentation conversion is performed on the three-dimensional voxel data, and the subsequent processing may be performed as three-dimensional voxel data.

[0081] Also, in this embodiment, the data storage unit 104 converts the data to be compressed into data obtained by normalizing and quantizing the coordinate information included in the data to be compressed and then compresses it. Therefore, while achieving a higher compression rate, it is possible to compress the identification information as it is, and thus it is possible to suppress the value of the identification information from being shifted due to quantization.

Embodiment

[0082] FIG. 13 is a diagram showing a configuration example of the data storage unit 104 according to the information compression system according to Embodiment 2 of the present disclosure. The data storage unit 104 shown in FIG. 13, as a configuration for compression processing, in addition to the configuration shown in FIG. 12, further includes an ID conversion compressor 301 and a voxel encoder 302, and as a configuration for decompression, in addition to the configuration shown in FIG. 12, further includes an ID conversion decompressor 311, and instead of the PC converter 213, further includes a voxel decoder / PC converter 213a.

[0083] In the compression process, the ID conversion compressor 301 exchanges the value of the identification information based on the similarity of the object and meaning indicated by the identification information. The ID conversion compressor 301 is constructed, for example, with a learned model using a DNN or the like.

[0084] FIG. 14 is a diagram for explaining the conversion process by the ID conversion compressor 301. As shown in FIG. 14(a), normally, since the "meaning" identified by the meaning ID is set independently of the value of the meaning ID, the distance between the values of the meaning ID and the similarity of the "meaning" (semantic distance) are independent. Therefore, if the meaning ID is compressed as it is, there is a risk that the "meaning" before and after decompression will change significantly due to the shift of the value of the meaning ID caused by compression.

[0085] Therefore, as shown in FIG. 14(b), the ID conversion compressor 301 converts the value of the meaning ID according to the meaning. For example, the ID conversion compressor 301 converts so that the values of the meaning IDs with a close semantic distance such as "car" and "road" are close.

[0086] Note that in FIG. 14, the value of the meaning ID before conversion is an integer value. The value of the meaning ID after conversion is not limited to an integer value. Also, the value of the meaning ID after conversion may have a width. For example, when the value of the meaning ID is from 0.5 to 1.1, that meaning ID may represent "car" as a meaning. Also, note that in FIG. 14, the meaning ID is used as an example, but the conversion compressor 301 may convert the instance ID in the same way as the meaning ID.

[0087] Returning to the description of FIG. 13. The voxel encoder 302 quantizes the voxel information generated by the voxelizer 202, and converts it into a feature map by encoding the quantized voxel information by an irreversible transformation. The voxel encoder 302 is constructed by a trained model using a DNN such as a CNN. Also, for the quantized voxel value, the value of the most common identification information among the identification information included in the quantization range is selected.

[0088] The entropy estimator 203 estimates the probability distribution as the entropy of the feature map generated by the voxel encoder 302, and the entropy encoder 204 encodes the feature map based on the probability distribution to generate compressed binary data.

[0089] In the stretching process, the entropy decoder 211 decodes the compressed binary data to generate a feature map. The voxel decoder / PC converter 213a decodes the feature map generated by the entropy decoder 211 to generate voxel information, and converts the voxel information into quantization chunk data having coordinate information and identification information.

[0090] The inverse quantization / renormalization unit 214 performs inverse quantization and renormalization on the coordinate information of the quantization chunk data, and the ID conversion expander 311 performs inverse conversion of the conversion by the ID conversion compressor 301 on the ID information of the quantization chunk data to generate chunk data as the expanded data.

[0091] As described above, according to this embodiment, since the value of the identification information is swapped based on semantic similarity and then the identification information is compressed, it is possible to suppress the occurrence of a semantic shift even when irreversible compression of the identification information is performed. Therefore, it is possible to further improve the compression ratio.

Embodiment

[0092] FIG. 15 is a diagram showing the configuration of the data storage unit 104 according to the information compression system according to Embodiment 3 of the present disclosure. The data storage unit 104 shown in FIG. 15 includes a semantic ID conversion compressor 401, an instance ID conversion compressor 402, and a point cloud compressor 403 as components for compression processing.

[0093] In this embodiment, the data storage unit 104 processes each chunk data of the data to be compressed as data (x, y, z, S, I) in a list format composed of components of coordinate information (x, y, z) and identification information (S, I).

[0094] The semantic ID conversion compressor 401 and the instance ID conversion compressor 402 constitute an ID color converter that converts the identification information (S, I) into color information in the form of (R, G, B). Thereby, the data in the list format (x, y, z, S, I) is converted into the data in the list format (x, y, z, R, G, B) using color information. Note that since the color information (R, G, B) included in the data in the list format (x, y, z, R, G, B) is converted from the identification information (S, I), unlike the color information included in the original color camera data, it is possible to reduce randomness and increase the compression ratio.

[0095] The point cloud compressor 403 generates compressed binary data obtained by compressing the information in the list format (x, y, z, R, G, B) as point cloud data with color information. As the point cloud compressor 403, an existing compressor for compressing point cloud data can be used.

[0096] Note that as a configuration for the decompression process, the data storage unit 104 has, for example, a point cloud decompressor that generates decompressed data obtained by decompressing the compressed binary data generated by the point cloud compressor 303, and a color ID converter that inverse-converts the identification information of the decompressed data generated by the point cloud decompressor to generate data in the list format (x, y, z, S, I) (both not shown).

[0097] Also, in the above description, the sensor data had point cloud data and color camera data, but it may not have point cloud data. In this case, for example, the data acquisition unit 101 may acquire point cloud data from the color camera data by analyzing the color camera data to estimate the position of the surface of the object, or may compress the sensor data without using the point cloud data.

[0098] FIG. 16 is a diagram showing an example of the configuration of the data storage unit 104 when compressing sensor data without using point cloud data. The data storage unit 104 shown in FIG. 16 has, as a configuration for the compression process, a semantic ID conversion compressor 411, an instance ID conversion compressor 412, and a moving image compressor 413.

[0099] In the example of FIG. 16, the data store unit 104 regards each chunk data of the data to be compressed as data of a three-dimensional array ([2][x][y]) indicating identification information (S, I), the pixel position in the vertical axis direction, and the pixel position in the vertical axis direction. Here, [2] indicates identification information, [X] indicates the pixel position in the vertical axis direction, and [Y] indicates the pixel position in the vertical axis direction.

[0100] The meaning ID conversion compressor 411 and the instance ID conversion compressor 412 constitute an ID color converter that converts the identification information (S, I) into information in the form of color information (R, G, B). As a result, the data of the three-dimensional array ([2][x][y]) is converted into the data of the three-dimensional array ([3][x][y]) using color information. Note that since the color information [3] included in the data of the three-dimensional array ([3][x][y]) is converted from the identification information (S, I), unlike the color information included in the original color camera data, it is possible to reduce randomness and increase the compression ratio.

[0101] The moving image compressor 413 compresses the information in the form of a three-dimensional array ([3][x][y]) as image data (more specifically, moving image data). As the moving image compressor 413, an existing compressor for compressing moving images can be used.

[0102] As a configuration for decompression processing, the data store unit 104 has, for example, a moving image decompressor that generates decompressed data obtained by decompressing the compressed binary data generated by the moving image compressor 413, and a color ID converter that inverse-converts the identification information of the decompressed data generated by the moving image decompressor to generate data of a three-dimensional array ([2][x][y]) (both not shown).

[0103] Note that in the case of the configuration of FIG. 16, the processing of the integration unit 103 is omitted. Also, in the case of the configuration of FIG. 16, the specification of the spatial range of the data to be decompressed is performed, for example, by specifying a sensor. Also, the specification of the spatial range of the data to be decompressed may be performed by specifying a position. In this case, based on the sensor information table, the sensor ID is specified from the specified position.

[0104] As described above, in this embodiment, the data to be compressed is converted into data in a list format including color information or data in a three-dimensional array and then compressed. Therefore, it is possible to improve the compression efficiency by using an existing compressor. Further, since point cloud data can be obtained from the image data, it is not necessary to use a sensor or the like for obtaining the point cloud data.

Example

[0105] FIG. 17 is a diagram showing the configurations of the segmentation unit 102 and the data store unit 104 according to the information compression system according to Example 4 of the present disclosure. In the example of FIG. 17, the sensor data does not have point cloud data. Further, the segmentation unit 102 and the data store unit 104 are integrated.

[0106] The segmentation unit 102 and the data store unit 104 include an encoder 501, an entropy estimator 502, and a decoder 503 as configurations for compression processing, and a generation decoder 504 as a configuration for expansion.

[0107] The encoder 501 encodes and compresses an input image ([3][x][y]), which is color camera data, by non-invertible transformation, thereby converting the input image (image[3][x][y]) into encoded data (z[c][a][b]). The encoded data is, for example, a feature map.

[0108] The entropy estimator 502 estimates the entropy (probability distribution) of the encoded data generated by the image encoder 501, and executes entropy encoding and entropy decoding processing in the same manner as in FIG. 12.

[0109] The decoder 503 decodes the encoded data generated by the image encoder 501 to generate output segmentation data (seg[2][x][y]), which is compressed binary data.

[0110] The encoder 501 and the decoder 503 are constructed using a learned model with a DNN such as a CNN. In order to construct a learned model by machine learning, for example, end-to-end learning combining the segmentation unit 102 and the data store unit 104 is used. As the loss function in learning, for example, "Loss = λ * entropy + distortion(seg data, teacher seg data)" is used. Here, λ is a parameter for determining the rate-distortion trade-off, and "entropy" is the entropy (information amount) calculated by the entropy estimator 502. Also, seg data is the output segmentation data, and teacher seg data is the teacher data of the output segmentation data. The distortion function is, for example, typically a general loss function in segmentation such as cross entropy, or mean squared error (MSE), etc. In addition, a differentiable image quality metric in the image, such as MS-SSIM (Multi-Scale Structural Similarity), etc. may be used.

[0111] The generation decoder 504 decodes the output segmentation data to generate output image data (image[3][x][y]) having color information as extended data. The generation decoder 504 is constructed using a model by a DNN learned by, for example, GAN (Generative Adversarial Network). The generation decoder 504 may be constructed by performing learning different from the learning for generating the encoder 501 and the decoder 503.

[0112] As described above, according to this embodiment, the segmentation unit 102 and the data store unit 104 are integrally constructed using a learned model. Therefore, it is possible to simplify the configuration of the information compression system. Also, since image data with color information is output in the extension process, data analysis, etc. can be performed using the same type of application program as before.

Example

[0113] FIG. 18 is a diagram showing the configuration of the data store unit 104 according to the information compression system of Example 5 of the present disclosure. The data store unit 104 shown in FIG. 18 has, as a configuration for compression processing, a normalizer 601, a point cloud encoder 602, an entropy estimator 603, and an entropy encoder 604, and has, as a configuration for decompression, an entropy decoder 611, an entropy estimator 612, a voxel decoder 613, and a renormalization / mesh generation unit 614.

[0114] In this example, for each piece of identification information, the coordinate information is compressed by constructing chunk data. Therefore, the identification information is held as one of the items of the chunk information 60.

[0115] The normalizer 601 normalizes the coordinate information. The point cloud encoder 602 encodes and quantizes the normalized coordinate information to generate a feature map. The feature map is a data array of a predetermined size, and the values are integerized by quantization. The point cloud encoder 602 is constructed by, for example, a DNN having a combination of an MLP (Multilayer perceptron) and a max pooling layer.

[0116] The entropy estimator 603 estimates a probability distribution as the entropy of the feature map generated by the point cloud encoder 602. The entropy estimator 603 is constructed by, for example, a trained model using a DNN such as a multi-layer 1D convolutional neural network.

[0117] The entropy encoder 604 encodes the feature map based on the probability distribution estimated by the entropy estimator 203 to generate compressed binary data.

[0118] Also, in the expansion process, the entropy decoder 611 decodes the compressed binary data to generate a feature map. Specifically, the entropy decoder 611 uses the entropy estimator 612 to generate a feature map. Specifically, the entropy decoder 211 uses the entropy estimator 612 to predict the probability distribution of the values (symbols) of the feature map, and the entropy decoder 611 decodes the symbols using the compressed binary data and the predicted probability distribution, and finally decodes them as a feature map.

[0119] The voxel decoder 613 decodes the feature map generated by the entropy decoder 611 for each three-dimensional region designated as a voxel to generate occupancy information indicating the occupancy of the object to be photographed in each region of the voxel. The occupancy, unlike the point cloud generally composed of only the position information of the surface of the photographed object, has a value close to 1 in the internal region of the photographed object and a value close to 0 in the external region of the photographed object. By using the voxel-shaped occupancy data, a mesh close to the original object can be obtained. Utilizing this characteristic and further using entropy encoding, it is possible to store the information of the photographed object with high accuracy with a smaller amount of data. The voxel decoder 613 is generated, for example, by a learned model using a DNN having an MLP.

[0120] The renormalization / mesh generation unit 614 generates coordinate information and surface information by renormalizing and meshing the occupancy information. The surface information indicates a surface represented by a set of three or more points indicated by the coordinate information. Meshing is performed, for example, by using the Marching cubes method or the like.

[0121] As described above, according to this embodiment, the coordinate information is compressed for each identification information. Also, in this embodiment, since the mesh information can be decoded from the compressed binary data, the mesh data can be efficiently read from the data utilization unit 106 without waste through the data conversion unit 105.

[0122] The above-described embodiments of the present disclosure are examples for the description of the present disclosure, and are not intended to limit the scope of the present disclosure only to those embodiments. Those skilled in the art can implement the present disclosure in various other ways without departing from the scope of the present disclosure.

Explanation of Signs

[0123] 1: Node 2: Sensor 3: Input / output device 101: Data acquisition unit 102: Segmentation unit 103: Integration unit 104: Data storage unit 105: Data conversion unit 106: Data utilization unit 201: Quantizer 202: Voxelizer 203: Entropy estimator 204: Entropy encoder 211: Entropy decoder 212: Entropy estimator 213, 213a: PCizer 214: Renormalizer 301: ID conversion compressor 302: Voxel encoder 303: Point cloud compressor 311: ID conversion expander 401: Semantic ID conversion compressor 402: Instance ID conversion compressor 403: Point cloud compressor 411: Semantic ID conversion compressor 412: Instance ID conversion compressor 413: Moving image compressor 501: Encoder 502: Entropy estimator 503: Decoder 504: Generation decoder 601: Normalizer 602: Point cloud encoder 603: Entropy estimator 604: Entropy encoder 611: Entropy decoder 612: Entropy estimator 613: Voxel decoder 614: Renormalization / mesh generation unit

Claims

1. An information compression system for compressing data, comprising: an acquisition unit configured to acquire the data; a generation unit configured to determine each object depicted in the data and the meaning of each object, and generate compression target data by converting the value of each element of the data into identification information representing each object and the meaning of each object based on the determination result; a compression unit configured to generate compressed data by compressing the compression target data; wherein the acquisition unit further acquires point cloud data having a plurality of coordinate information indicating each position on the surface of the object; and the generation unit generates integrated data having the coordinate information and the identification information representing the object at the position for each position on the surface as the compression target data based on the point cloud data and the determination result.

2. The information compression system according to claim 1, wherein the acquisition unit acquires the point cloud data from the data.

3. The information compression system according to claim 1, wherein the compression unit compresses the compression target data by converting the compression target data into data obtained by normalizing and quantizing the coordinate information included in the compression target data.

4. The information compression system according to claim 1, wherein the compression unit compresses the compression target data by converting the compression target data into data obtained by swapping the values of the respective identification information included in the compression target data based on the similarity between the objects and the meanings.

5. The information compression system according to claim 1, wherein the compression unit compresses the compression target data by converting the compression target data into data in a list format in which the value of the coordinate information and the value representing a color generated from each identification information included in the compression target data are arranged side by side.

6. The information compression system according to claim 1, wherein the compression unit compresses the compression target data by converting the compression target data into three-dimensional array format data having the values of the identification information as the values of each element.

7. The information compression system according to claim 1, wherein the generation unit and the compression unit are integrated and constructed using a learned model.

8. The information compression system according to claim 1, wherein the compression unit compresses the coordinate information included in the compression target data for each identification information.

9. The information compression system according to claim 1, further comprising: a decompression unit configured to generate decompressed data by decompressing the compressed data; and a generation unit configured to process and output the decompressed data.

10. further comprising a data utilization unit that provides an interface for specifying at least one of the space and time at which the data was acquired, The information compression system according to claim 9, wherein the expansion unit expands the compressed data according to the specified content specified by the interface.

11. A compression method by an information compression system that compresses data, comprising: acquiring the data, determining each object appearing in the data and the meaning of each object, and generating compression target data in which the value of each element of the data is converted into identification information representing each object and the meaning of each object based on the determination result, generating compressed data obtained by compressing the compression target data, In the acquisition of the data, point cloud data having a plurality of coordinate information indicating each position on the surface of the object is further acquired, In the generation of the compression target data, integrated data having the coordinate information and the identification information representing the object at the position for each position on the surface is generated as the compression target data based on the point cloud data and the determination result. Information compression method.

Citation Information

Patent Citations

  • Information conversion system

    JP2000295611A

  • Storage system and storage control method

    JP2021111882A

Cited By

  • Method for processing data sets containing at least one time series, device for carrying out, vehicle and computer program

    US12738957B2

  • Method for processing data sets containing at least one time series, device for carrying out, vehicle and computer program

    US20230111292A1