Diversity degree concentration device, diversity degree concentration method and diversity degree concentration program

The diversity enrichment device and method address the challenge of maintaining data group diversity by identifying data for exclusion based on a specific distribution, optimizing retention intervals, and adapting to operational changes.

JP2025096738AActive Publication Date: 2025-06-30早坂壮大
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023212621
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-18
Publication Date
2025-06-30
Estimated Expiration
2043-12-18

AI Technical Summary

Technical Problem

Existing techniques for updating data groups struggle to maintain diversity, particularly in terms of data age, due to limitations in setting retention intervals and optimizing parameter settings based on changing operational conditions.

Method used

A diversity enrichment device and method that identifies data to be excluded from a data group by specifying data whose exclusion would approximate a specific distribution of attribute values, such as an ideal distribution composed of proportional and exponential phases.

Benefits of technology

This approach allows for efficient maintenance of data group diversity by suitably identifying data with poor contribution to diversity for exclusion, thereby optimizing the retention interval and adapting to changes in operational status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025096738000001_ABST
    Figure 2025096738000001_ABST
Patent Text Reader

Abstract

To provide a diversity degree concentration device, method and program which can suitably identify data to be eliminated from a group of data and efficiently maintain diversity of the group of data.SOLUTION: A diversity degree concentration device comprises: identification means which identifies data to be eliminated from among the group of data which is a collection of data including attribute values so that distribution of the attribute values in the group of data after elimination of the data to be eliminated resembles specific distribution; and output means which outputs information indicating the data to be eliminated identified by the identification means.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for updating a data group.

Background Art

[0002] When adding new data to a group of related data such as a data group obtained by experiments or observations or a backup data group obtained by modifying original data, in order to avoid straining the storage capacity of the storage medium, some data may be deleted (i.e., thinned out).

[0003] As a technique applied when deleting some data from a group of data, for example, there is a technique of deleting data in order from the data with an old generation time. Such a technique is generally based on the presumption that older data is relatively less useful. However, when applying such a technique, data before a certain point in time is lost, so the diversity such as data age of the entire group of data is impaired.

[0004] A technique of also making new data with a new generation time a deletion target and retaining old data at a predetermined ratio is based on the presumption that a certain degree of usefulness exists in old data. One of the important things when applying this type of technique to update a group of data is to thin out the data without impairing the diversity such as data age of the entire group of data as much as possible.

[0005] As a related technique, Non-Patent Document 1 describes retaining a group of data so that the data age is non-uniform. Specifically, on page 14 of Non-Patent Document 1, a technique is described in which for each period of from 1 day ago to 7 days ago, from 8 days ago to 90 days ago, from 91 days ago to 365 days ago, and 366 days or later for the generations saved by backup, the intervals for retaining the file versions can be set respectively.

Prior Art Documents

Non-Patent Documents

[0006] [Non-Patent Document 1] "Welcome to Getting Started BackStore by PROe! Operation Manual Ver 5.0.1", [online], Nekojarashi Co., Ltd., [searched on September 28, Reiwa 5], Internet <URL:https: / / www.backstore.jp / dl / BackStore_GettingStarted_PROe.pdf> [Summary of the Invention] [Problems to be Solved by the Invention]

[0007] However, in the technology described in Non-Patent Document 1, the interval for retaining file versions can only be set at fixed intervals determined in advance. That is, in the technology described in Non-Patent Document 1, the period during which the interval for retaining file versions can be set is limited. Also, in the technology described in Non-Patent Document 1, since it is necessary to perform a plurality of parameter settings in advance, it is difficult to optimize the retention interval according to changes in the operation status, etc. over time. Therefore, it is difficult to efficiently maintain the diversity of a group of data in the technology described in Non-Patent Document 1.

[0008] An object of the present invention is to provide a diversity enrichment device, a diversity enrichment method, and a diversity enrichment program that can suitably identify data to be excluded from a group of data and efficiently maintain the diversity of the group of data. [Means for Solving the Problems]

[0009] The diversity enrichment device according to the present invention includes: a specifying means for specifying data to be excluded from a group of data, which is a set of data including attribute values, such that the distribution of the attribute values in the group of data after the data to be excluded is excluded approximates a specific distribution; and an output means for outputting information indicating the data to be excluded specified by the specifying means.

[0010] The diversity enrichment method according to the present invention is characterized in that a computer identifies data to be excluded from a group of data, which is a set of data including attribute values, so that the distribution of the attribute values in the group of data after the data to be excluded is excluded approximates a specific distribution, and outputs information indicating the identified data to be excluded.

[0011] The diversity enrichment program according to the present invention is characterized in that it causes a computer to perform a specific process of identifying data to be excluded from a group of data, which is a set of data including attribute values, so that the distribution of the attribute values in the group of data after the data to be excluded is excluded approximates a specific distribution, and an output process of outputting information indicating the data to be excluded identified by the specific process.

Advantages of the Invention

[0012] According to the present invention, it is possible to suitably identify data with poor contribution to diversity as data to be excluded from a group of data, and efficiently maintain the diversity of the group of data.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Embodiments for Carrying Out the Invention

[0014] Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0015] FIG. 1 is a block diagram showing an example of a system to which a diversity concentration device is applied. The diversity concentration system to which the diversity concentration device illustrated in FIG. 1 is applied includes a storage unit 1, a monitoring means 2, an extraction means 3, an exclusion means 4, a diversity concentration device 10, and a user terminal 20. The diversity concentration device 10 includes an input means 5, an exclusion target specifying means 6, and an output means 7. Further, the exclusion target specifying means 6 includes an ideal distribution specifying means 61 and a post-change distribution evaluation means 62. Note that the arrows in FIG. 1 simply indicate the direction of the flow of signals (data) and do not exclude bidirectionality.

[0016] The storage unit 1 is realized by a storage device such as a semiconductor memory or a hard disk, for example. Each function of the monitoring means 2, the extraction means 3, and the exclusion means 4 is realized, for example, by a CPU (Central Processing Unit) of an information processing device executing processing according to a program stored in a storage medium. Further, the input means 5, the exclusion target specifying means 6 (including the ideal distribution specifying means 61 and the post-change distribution evaluation means), and the output means are realized, for example, by a CPU of the diversity concentration device 10 executing processing according to a program stored in a storage medium.

[0017] The monitoring means 2, the extraction means 3, and the elimination means 4 may be included in, for example, the same information processing apparatus, or may be included in different information processing apparatuses respectively. Further, the monitoring means 2, the extraction means 3, and the elimination means 4 may be included in, for example, an information processing apparatus (not shown) including the storage unit 1, or may be included in the diversity concentration apparatus 10. Further, the storage unit 1, the monitoring means 2, the extraction means 3, and the elimination means 4 may be included in the diversity concentration apparatus 10.

[0018] The storage unit 1 stores a group of data to be managed. The storage unit 1 may store various data related to the group of data. For example, an information processing apparatus (not shown) inside or outside the diversity concentration system can store new data in the storage unit 1 as a group of data. Further, for example, an information processing apparatus (not shown) inside or outside the diversity concentration system can update the data stored in the storage unit 1 as a group of data.

[0019] A group of data to be managed is a set of data each including one or more attribute values. The attribute value may have one or more unique reference values and may have the feature that the distance between the attribute value and the reference value can be compared.

[0020] A group of data may be a set of the data itself, or may be a set of metadata indicating attributes, characteristics, structures, meanings, relationships, etc. of the data. Further, a group of data may be a set of data (hereinafter also referred to as catalog data) indicating attributes, characteristics, structures, meanings, relationships, etc. of objects or events not expressed by data. Further, different types of data such as text data, spreadsheet data, video data, and audio data may be mixed in a group of data.

[0021] In this embodiment, the expression "data including attribute values" is used. However, the attribute values included in the data may indicate the attributes of the data itself, or may indicate the attributes of other data. Further, the attribute values included in the data may indicate the attributes of objects or events that are not represented by data.

[0022] The attribute values included in the data may be, for example, values indicating the date and time associated with the data. Specifically, the attribute values included in the data may be determined based on the date and time when access was made to the data, or the date and time when the data was generated.

[0023] In this embodiment, each piece of data included in a group of data to be managed is stored in association with sequence information (hereinafter also referred to as rank or Rank) for identifying the order in which the data is stored in the group of data. Note that, not limited to the configuration of this embodiment, for example, the sequence information corresponding to the data may be configured to be generated based on the attribute values of the data as needed.

[0024] Next, a specific example of a group of data to be managed will be described. For example, a group of data D is composed of c text files D1 to Dc. Each piece of data (i.e., text file) in the group of data D has, as an attribute value, information indicating the time T when the file was last accessed. The group of data D is a set of document data generated each time a specific document data is modified. When each text file is plotted on a one-dimensional time axis based on the attribute value (time T), the distribution state of each text file (i.e., the attribute value (time T)) is related to the diversity of the last access times of the group of data D. For example, when the distribution state of the text files (i.e., the attribute values (time T)) is close to an ideal distribution state, the group of data D is determined to have high diversity. Conversely, when the distribution state of the text files (i.e., the attribute values (time T)) is far from the ideal distribution state, the group of data D is determined to have low diversity.

[0025] In a group of data D, the text file D1 is the most recently accessed (i.e., the newest) text file. On the other hand, the text file Dc is the oldest text file. In this case, rank 1 with a rank value of 1 is associated with the text file D1 as sequence information. Also, rank c with a rank value of c is associated with the text file Dc as sequence information. That is, in this embodiment, it is shown that the larger the rank value, the older the data and the longer the period it has been stored as a group of data. Further, in this embodiment, for example, when new data is stored in a group of data, rank 1 is associated with the new data, and the rank value of each of the already stored data is incremented by 1. With such a configuration, in this embodiment, the order in which each of the group of data to be managed is stored can be identified. Hereinafter, the rank with the smallest rank value (i.e., associated with the newest data) is also referred to as the minimum rank or the latest rank. Also, the rank with the largest rank value (i.e., associated with the oldest data) is also referred to as the maximum rank or the oldest rank.

[0026] In this embodiment, for the sake of simplicity of explanation, a configuration in which data such as a single text file is stored in association with a single rank will be described. However, for example, a configuration may be adopted in which a plurality of data having common storage timings, attribute values, etc. are stored in association with a single rank.

[0027] In this embodiment, each piece of data included in a group of data is stored in association with a group value that enables identification of whether it belongs to any of a plurality of types of groups. In this embodiment, as groups of data, a first group that is a monitoring target and an exclusion target, and a second group that is a monitoring target but not an exclusion target are provided. As a group of data, for example, a third group that is not a monitoring target may be provided. Data belonging to the first group is stored in association with group value = 1. Data belonging to the second group is stored in association with group value = 2. Data belonging to the third group is stored in association with group value = 3. The group value of each piece of data is specified, for example, based on a specifying operation performed by the user using the user terminal 20. The group value of each piece of data may be automatically specified according to a predetermined algorithm, for example, based on the attributes, characteristics, structure, meaning, and relationship of the data.

[0028] The monitoring means 2 has a function of monitoring a group of data to be managed. The monitoring means 2 monitors, for example, whether it is necessary to exclude some data from a group of data to be managed. Specifically, the monitoring means 2 monitors the data amounts of the data with group value = 1 and the data with group value = 2 among the group of data stored in the storage unit 1. When the monitoring means 2 determines that the data amount exceeds a predetermined capacity due to the occurrence of addition of new data to the group of data or the like, it notifies the extraction means 3 to that effect.

[0029] In this embodiment, for the sake of simplifying the explanation, the case where the upper limit of the number of ranks for a group of data to be managed is fixedly determined in advance will be described. In this case, when the total number of data with group value = 1 or 2 (for example, the total number of the above-described text files) among the group of data to be managed exceeds the predetermined upper limit number of ranks (for example, 100), the monitoring means 2 notifies the extraction means 3 to that effect.

[0030] In this embodiment, the upper limit of the number of ranks for a group of data to be managed (i.e., the total upper limit of data) is fixedly determined in advance. However, the upper limit of the number of ranks may be variable. For example, the number of ranks for a group of data to be managed may be fixedly determined according to the amount of resources available for storage (hereinafter also referred to as the supply amount), or may be variable according to changes in the supply amount.

[0031] The supply amount may be determined, for example, based on the total amount or total number of data that can be stored in a specified storage medium. Also, the supply amount may be determined, for example, based on the amount of communication communicated in a predetermined period by communication means (not shown) included in the diversity concentration device 10. Further, the supply amount may be determined based on the amount of calculation calculated in a predetermined period by calculation means (not shown) such as a CPU included in the diversity concentration device 10. Also, the supply amount may be determined, for example, based on a value input by a user via user input means (not shown) such as a keyboard.

[0032] The extraction means 3 has a function of extracting data from the storage unit 1. The extraction means 3 extracts, for example, data with a group value = 1 or 2 from a group of data to be managed based on a notification from the monitoring means 2. Also, the extraction means 3 transmits the extracted data to the diversity concentration device 10. Note that the extraction means 3 may transmit only the metadata of the data instead of the data itself with a group value of 1 or 2 to the diversity concentration device 10. When a group of data is a set of directory data, the extraction means 3 transmits the directory data with a group value of 1 or 2 to the diversity concentration device 10.

[0033] In this embodiment, the extraction means 3 is configured to perform data extraction processing and transmission processing based on the notification from the monitoring means 2. However, for example, the extraction means 3 may perform data extraction processing and transmission processing based on an instruction operation performed by the user using the user terminal 20. That is, a configuration may be adopted in which an instruction can be given to exclude some data from a group of data according to the user's judgment. For example, the user may use the user terminal 20 to input extraction instruction information indicating the range and conditions of the data to be extracted from a group of data, and exclusion target information indicating the amount of data to be excluded. In this case, the extraction means 3 transmits the exclusion target information to the diversity concentration device 10 together with the data extracted based on the extraction instruction information. Further, the diversity concentration device 10 repeats the process of identifying the data to be excluded until the exclusion target is achieved.

[0034] The input means 5 of the diversity concentration device 10 has a function of inputting data. The input means 5 inputs, for example, a group of data transmitted by the extraction means 3.

[0035] The exclusion target identification means 6 has a function of identifying the data to be excluded from a group of data input by the input means 5. The exclusion target identification means 6 identifies, for example, the data to be excluded from the data with a group value = 1 among the group of data input by the input means 5.

[0036] The ideal distribution identification means 61 included in the exclusion target identification means 6 has a function of identifying an ideal distribution indicating an ideal distribution state with high diversity regarding the distribution of the attribute values of a group of data. Note that the ideal distribution state with high diversity is not uniquely determined, and any desirable distribution state may be used. The ideal distribution may be determined, for example, based on information indicating conditions input by the user via a user input means such as a keyboard.

[0037] In this embodiment, various distributions can be applied as the ideal distribution. For example, the ideal distribution may be composed of a single phase in which the distribution is specified by one function, or may be composed of a plurality of phases in which the distributions are respectively specified by different functions.

[0038] For example, the ideal distribution when the upper limit total number of data stored as a group of data is 9 will be described. In this case, when the range of the attribute value is 1 to 260 and the minimum increment of the attribute value is 1, if an exponential distribution is applied to the entire region as the ideal distribution, a desirable distribution state can be efficiently maintained, such as {1, 2, 4, 8, 16, 32, 64, 128, 256}. However, for example, when the range of the attribute value is 1 to 14 and the minimum increment of the attribute value is 1, this is not the case. Specifically, the ideal distribution will include a small number like 1.1, but since the significant increment is 1, rounding to the nearest whole number will result in duplicates such as {1, 1, 2, 3, 4, 5, 6, 10, 14}. Thus, when the range with respect to the increment of the attribute value is small, even if an exponential distribution is applied to the entire region as the ideal distribution, a desirable distribution state cannot be efficiently maintained.

[0039] In the present embodiment, in order to efficiently maintain a desirable distribution state, an ideal distribution composed of a proportional phase to which a proportional distribution is applied and an exponential phase to which an exponential distribution is applied is applied. Hereinafter, the ideal distribution composed of the proportional phase and the exponential phase will be described.

[0040] FIG. 2 is an explanatory diagram showing an example of the ideal distribution composed of the proportional phase and the exponential phase. In FIG. 2, FIG. 2(1) shows the holding state of a group of data at the first time point. FIG. 2(2) shows the holding state of a group of data at the second time point after a predetermined period has elapsed from the first time point shown in FIG. 2(1). In the example shown in FIGS. 2(1) and 2(2), a process of excluding (i.e., thinning out) data from a group of data and adding new data (hereinafter, also referred to as an update process for a group of data) is performed a plurality of times during the predetermined period from the first time point to the second time point.

[0041] The horizontal axis (x-axis) of the graph shown in FIGS. 2(1) and 2(2) represents the rank value of the data as Rank. The vertical axis (y-axis) of the graph represents, on a logarithmic scale, the data age (e.g., the number of days or hours from the generation date and time to the current date and time) specified as the attribute value of the data as Age. The graph shown in FIGS. 2(1) and 2(2) is a phase diagram showing the relationship between Rank and Age. In this example, the rank k of a group of data is set as x k and the logarithm of Age corresponding to the rank k is set as y k . Also, assume that a group of data contains c pieces of data. Therefore, k takes values from 1 to c. As shown in FIGS. 2(1) and 2(2), the coordinates of the data with rank c corresponding to the oldest rank are (x c , y c ). Hereinafter, the coordinates (x c , y c ) of the data with rank c are also referred to as the oldest point c.

[0042] In each graph shown in FIGS. 2(1) and 2(2), the region above the transition curve corresponds to the proportional phase, and the region below the transition curve corresponds to the exponential phase. In this embodiment, various methods can be applied to determine the point (hereinafter also referred to as the transition point t) at which the transition occurs from the proportional phase to the exponential phase. For example, the method shown below may be applied.

[0043] As shown in FIGS. 2(1) and 2(2), the transition point t is expressed as (x t , y t ). In this case, the ideal distribution in the proportional phase (x ≤ x t ) and the ideal distribution in the exponential phase (x t < x) are smoothly connected at the transition point t. Also, the ideal distribution in the exponential phase passes through the transition point t and the oldest point c.

[0044] In the semi-logarithmic graph shown in FIG. 2, the ideal distribution in the proportional phase is represented by a logarithmic distribution. Also, the ideal distribution in the proportional phase is expressed by the following formula (1). In the semi-logarithmic graph shown in FIG. 2, the ideal distribution in the exponential phase is represented by a linear distribution. Also, since the ideal distribution in the exponential phase is a straight line passing through the oldest point c and the transition point t, it is expressed by the following formula (2). Since the slope of this straight line is equal to the slope at the transition point t, the transition point t will exist on the transition curve of the following formula (3). Therefore, the transition point t can be specified by identifying the intersection of the line connecting the transition curve and each data point of a group of data. When the intersection is not uniquely determined, the point closest to the oldest point c among the multiple intersections may be used as the transition point t.

[0045] [Number]

[0046] [Number]

[0047] [Number]

[0048] As shown in FIGS. 2(1) and (2), the position of the transition point t changes depending on the holding state of a group of data. That is, the position of the transition point t changes when an update process is performed on a group of data. Therefore, the proportional phase and the exponential phase also change in range when an update process is performed on a group of data.

[0049] Also, as shown in FIGS. 2(1) and (2), the ideal distribution changes depending on the holding state of a group of data. That is, the ideal distribution shown in FIGS. 2(1) and (2) changes when an update process is performed on a group of data. Therefore, by using the ideal distribution as shown in FIG. 2, it is possible to automatically optimize the distribution state such as the data holding interval according to changes in the operation status etc. over time.

[0050] The transition point t and the ideal distribution are obtained based on a group of data. In the present embodiment, the upper limit total number of data defined for a group of data corresponds to the value of the oldest rank (x c ). That is, when the upper limit total number of data defined for a group of data is 100, the value of the oldest rank (x c ) is 100. Therefore, in the present embodiment, the transition point t is based on the upper limit total number of data defined for a group of data, that is, the value of the oldest rank (x c ) specified by the upper limit total number, and the attribute value (y c ) of the data corresponding to the oldest rank, and is obtained. Further, the ideal distribution is obtained based on the transition point t and the oldest point c obtained in this way.

[0051] The ideal distribution is preferably specified by the outer edge obtained by connecting each data point in the attribute value space. For example, in a group of data on a one-dimensional age space, the oldest data and the latest data (or 0) form the outer edge. In the present embodiment, the state of being evenly dispersed on a logarithmic scale is defined as the ideal distribution. In a state where all data points are the outer edge (for example, likely to occur when the dimension of the attribute value is high relative to the number of data points), the ideal distribution may be specified by the degree of deviation from the center of gravity or median of the data points. For example, a state of being dispersed in a standard distribution with respect to the average point vector in a five-subject score space may be defined as the ideal distribution. Also, when data points can be clustered, this may be done for each cluster. For example, when there are science clusters and liberal arts clusters, a more realistic ideal distribution can be specified by defining the distribution in the same way for each and then adding them together.

[0052] When applying the ideal distribution composed of the proportional phase and the exponential phase shown in FIG. 2, the ideal distribution specifying means 61 first specifies the transition point t. Next, the ideal distribution specifying means 61 specifies the ideal distribution based on the specified transition point t.

[0053] In this embodiment, the ideal distribution is composed of a plurality of phases (i.e., a proportional phase and an exponential phase) with different distribution characteristics. With such a configuration, a desirable distribution state can be suitably expressed.

[0054] In this embodiment, the ideal distribution specifying means 61 specifies an ideal distribution composed of two phases, namely a proportional phase and an exponential phase. However, for example, the ideal distribution specifying means 61 may specify an ideal distribution composed of three or more phases with different distribution characteristics. Also, in this case, the ideal distribution specifying means 61 may be configured to determine whether to target the data belonging to each phase for exclusion.

[0055] Also, the ideal distribution specifying means 61 may specify the ideal distribution by, for example, using a histogram. Further, when a group of data is represented as a two-dimensional distribution, for example, the ideal distribution specifying means 61 may specify an approximation curve by the least squares method and use the specified approximation curve as the ideal distribution.

[0056] The post-change distribution evaluation means 62 included in the exclusion target specifying means 6 has a function of specifying the distribution of the attribute values of a group of data after a specific data is excluded from the group of data input by the input means 5 (hereinafter also referred to as the post-change distribution). Note that, for example, the distribution of the attribute values of each data when a specific data is excluded from a group of data and the latest data is added to the group of data may be used as the post-change distribution.

[0057] Generally, new data has a higher probability of being more useful than old data. Also, even for old data, if the oldest data is excluded, the diversity of the entire group of data may decrease. Therefore, in this embodiment, new data stored after the rank corresponding to the transition point t (hereinafter also referred to as the transition rank) and the oldest data corresponding to the oldest rank are configured not to be targets for exclusion. Thus, the post-change distribution evaluation means 62 specifies the post-change distribution when each data (excluding new data stored after the transition rank (i.e., data included in the proportional phase) and the oldest data corresponding to the oldest rank) is excluded from the group of data among the data that can be excluded (i.e., data with group value = 1). Note that, not limited to the configuration of this embodiment, new data stored after the transition rank (i.e., data included in the proportional phase) and the oldest data corresponding to the oldest rank may also be configured to be eligible for exclusion.

[0058] The post-change distribution evaluation means 62 has a function of evaluating the degree of disagreement (or degree of agreement) with the ideal distribution for each specified post-change distribution.

[0059] The post-change distribution evaluation means 62, for example, obtains the difference between the ideal distribution and the post-change distribution as a deviation. Next, the post-change distribution evaluation means 62 evaluates that the degree of disagreement is high (or the degree of agreement is low) when the deviation is large, and evaluates that the degree of disagreement is small (or the degree of agreement is high) when the deviation is small. Note that the evaluation of the degree of disagreement (or degree of agreement) between the ideal distribution and the post-change distribution is not limited to the exemplified method, and various methods can be applied.

[0060] The exclusion target specifying means 6 compares the degrees of inconsistency (or consistency) evaluated for each of the post-change distributions, and specifies the post-change distribution with the lowest degree of inconsistency (or the highest degree of consistency). That is, the exclusion target specifying means 6 specifies the post-change distribution that most closely approximates the ideal distribution from among a plurality of post-change distributions. Further, the exclusion target specifying means 6 specifies the data excluded when generating the specified post-change distribution as the data to be excluded. That is, the exclusion target specifying means 6 specifies the data to be excluded such that the distribution of a group of data when the data to be excluded is excluded most closely approximates the ideal distribution.

[0061] An example will be conceptually described with reference to FIG. 3 regarding the specification of the data to be excluded. FIG. 3 is an explanatory diagram showing an example of the specification of the data to be excluded. In the graph shown in FIG. 3, for each rank other than the oldest rank included in the exponential phase, the degree to which the post-change distribution when the data of that rank is excluded approaches the ideal distribution is shown as Gain. In the graph shown in FIG. 3, the scale on the right side of the graph corresponds to the degree of Gain. The value of Gain can be obtained, for example, by the exclusion target specifying means 6 specifying the post-change distribution when the corresponding data is excluded for each rank and evaluating the degree of inconsistency between the post-change distribution and the ideal distribution.

[0062] In the graph shown in FIG. 3, the data of the rank with the maximum Gain value is specified as the data to be excluded. It can be said that the data of the rank with the maximum Gain value is the data that most closely approximates the distribution of the attribute values of the group of data when the data is excluded from the group of data. In other words, it can also be said that the data of the rank with the maximum Gain value is the data that contributes little to the diversity of the group of data.

[0063] In this embodiment, the post-change distribution evaluation means 62 does not specify the post-change distribution when the data that cannot be excluded (i.e., the data with group value = 2) is excluded from each group of data. However, the post-change distribution evaluation means 62 may, for example, specify the post-change distribution when the exponential-phase data excluding the oldest-rank data is excluded from each group of data regardless of the group value of the data. In this case, after the exclusion target specifying means 6 specifies the post-change distribution that most closely approximates the ideal distribution, it determines whether the data excluded when generating the post-change distribution is data with group value = 1 (i.e., data that can be excluded). When it is data with group value = 1, the exclusion target specifying means 6 sets the data as an exclusion target. When it is not data with group value = 1, the exclusion target specifying means 6 then specifies the post-change distribution that approximates the ideal distribution. Next, the exclusion target specifying means 6 determines whether the data excluded when generating the specified post-change distribution is data with group value = 1 (i.e., data that can be excluded). The exclusion target specifying means 6 repeats the process of specifying the post-change distribution that approximates the ideal distribution and the process of determining whether it is data with group value = 1 until data with group value 1 is specified.

[0064] The output means 7 has a function of outputting information indicating the data to be excluded specified by the exclusion target specifying means 6.

[0065] The output means 7 outputs, for example, information indicating the data to be excluded to the user terminal 20. In this case, the user can use the user terminal 20 to confirm the specified data to be excluded. Further, the user can use the user terminal 20 to indicate whether to exclude the specified data to be excluded. Also, when a plurality of exclusion methods are provided, the user can use the user terminal 20 to indicate the exclusion method. When the user performs an input operation instructing exclusion execution using the user terminal 20, the user terminal 20 transmits information indicating the data to be excluded, information instructing exclusion execution, information specifying the exclusion method, etc. to the exclusion means 4.

[0066] The output means 7 may be configured to output information indicating the data to be excluded to the exclusion means 4, for example, without passing through the user terminal 20. That is, a process of excluding (i.e., thinning out) data from a group of data may be configured to be automatically performed. In this case, the output means 7 may also output information specifying an exclusion method (for example, a predetermined exclusion method, or an exclusion method automatically determined according to a predetermined algorithm based on the attributes, characteristics, structure, meaning, and relationship of the data to be excluded). Further, for example, the output means 7 may be configured to output the data to be excluded with the group value updated to a value other than 1 (for example, group value = 3) to the storage unit 1 without passing through the user terminal 20 and the exclusion means 4, so as to update a group of data.

[0067] The user terminal 20 is realized by an information processing device such as a personal computer, a smartphone, a tablet computer, etc., for example. For example, the user terminal 20 displays information indicating the data to be excluded output from the output means 7 on a display unit. Further, for example, the user terminal 20 transmits information indicating the data to be excluded, information instructing the execution of exclusion, information specifying an exclusion method, etc. to the exclusion means 4 based on an input operation by the user.

[0068] The exclusion means 4 has a function of excluding the data to be excluded from a group of data stored in the storage unit 1.

[0069] In the present embodiment, the expression of excluding data from a group of data is used, which includes changing the management form of the data in addition to deleting the data. For example, for specific data among a group of data, processes such as changing the group value to a value other than 1 (for example, group value = 3), migrating the storage destination to another storage medium, performing reversible compression or irreversible compression, and reducing the resolution are included in the process of excluding specific data from a group of data.

[0070] A case where moving image data composed of frame image data corresponding to each of a series of consecutive frames is held as a series of data will be described as an example of a process of excluding data from a series of data. In this case, the process of excluding data from a series of data includes, in addition to the process of deleting one frame image data specified as an exclusion target, the process of reducing the resolution of the one frame image data, and the process of reversibly or irreversibly compressing the one frame image data.

[0071] The exclusion means 4 excludes the data to be excluded based on, for example, the information output from the output means 7 or the information transmitted from the user terminal 20. Further, the exclusion means 4 excludes the data to be excluded by, for example, a preset exclusion method or an exclusion method specified by the information transmitted from the output means 7 or the user terminal 20.

[0072] Next, the processing operation of the diversity concentration device 10 will be described. FIG. 4 is a flowchart showing an example of the processing operation of the diversity concentration device 10.

[0073] When the extraction means 3 transmits a series of data extracted to the diversity concentration device 10, the input means 5 receives and inputs the series of data (Y in step S1).

[0074] Next, the ideal distribution specifying means 61 specifies the transition point t between the proportional phase and the exponential phase in the ideal distribution corresponding to the series of data input by the input means 5 (step S2).

[0075] Next, the ideal distribution specifying means 61 specifies the ideal distribution based on the specified transition point t (step S3).

[0076] Next, the post-change distribution evaluation means 62 substitutes the value obtained by subtracting 1 from the value of the oldest rank c into the variable k (step S4). That is, the post-change distribution evaluation means 62 specifies the rank k corresponding to the data one newer than the oldest rank c.

[0077] Next, the post-change distribution evaluation means 62 determines whether the data of rank k is data that can be excluded (i.e., data with group value = 1) (step S5). If it is not data that can be excluded, the process proceeds to step S8.

[0078] If the data of rank k is data that can be excluded, the post-change distribution evaluation means 62 specifies the post-change distribution when the data corresponding to rank k is excluded from a group of data (step S6).

[0079] Next, the post-change distribution evaluation means 62 evaluates the degree of mismatch between the specified post-change distribution and the ideal distribution (step S7).

[0080] Next, the post-change distribution evaluation means 62 subtracts 1 from the value of variable k (step S8). Then, the post-change distribution evaluation means 62 determines whether the value of variable k matches the value of the transition rank t (step S9).

[0081] If the value of variable k does not match the value of the transition rank t (N in step S9), the process proceeds to step S5. By repeatedly executing the processes of steps S5 to S9, the post-change distribution when data is excluded from the data corresponding to the rank that is one newer than the oldest rank c to the data corresponding to the rank that is one older than the transition rank t is specified.

[0082] When the value of variable k matches the value of the transition rank t (Y in step S9), the exclusion target specifying means 6 compares the degrees of mismatch evaluated for each post-change distribution and specifies the post-change distribution with the lowest degree of mismatch (step S10).

[0083] Next, the exclusion target specifying means 6 specifies the data excluded when generating the specified post-change distribution as the data to be excluded (step S11).

[0084] Next, the output means 7 outputs information indicating the data to be excluded (step S12).

[0085] In the example shown in FIG. 4, the diversity concentration device 10 ends the process based on outputting information indicating the data to be excluded that has been specified. However, for example, the diversity concentration device 10 may be configured to end the process when an end condition such as achieving the exclusion target (that is, specifying the data to be excluded among the amount of data specified as the exclusion target) is satisfied. That is, the diversity concentration device 10 may be configured to continue the process of specifying the data to be excluded until the end condition is satisfied. In this case, for example, the diversity concentration device 10 may output information indicating the data each time the data to be excluded is specified. Further, for example, the diversity concentration device 10 may output the information indicating the data in a lump after specifying a plurality of data to be excluded that achieve the exclusion target.

[0086] In the present embodiment, the exclusion target specifying means 6 specifies the data to be excluded from a group of data (for example, a set of data itself, a set of metadata, a set of directory data, etc.) which is a set of data including attribute values so that the distribution of the attribute values in the group of data after the data to be excluded is excluded approximates the ideal distribution. With such a configuration, in the present embodiment, data with little contribution to diversity can be preferably specified as the data to be excluded from a group of data, and the diversity of the group of data can be efficiently maintained.

[0087] In the present embodiment, the ideal distribution specifying means 61 specifies the ideal distribution based on a predetermined total amount (for example, the upper limit total number of a group of data. The oldest point c is specified by the upper limit total number) defined for a group of data. Further, the post-change distribution evaluation means 62 specifies the post-change distribution which is the distribution of the attribute values when any one data is excluded from a group of data. Next, the exclusion target specifying means 6 specifies the data excluded when generating the post-change distribution with the highest degree of coincidence with the ideal distribution or the post-change distribution with the lowest degree of non-coincidence with the ideal distribution as the data to be excluded. With such a configuration, in the present embodiment, it is possible to specify the data to be excluded so as to approximate the most desirable distribution state of the attribute values.

[0088] Next, the differences between the case of managing backup data by applying a backup method of deleting from old data and the case of managing backup data by applying the diversity concentration method of the present embodiment using the diversity concentration device 10 will be described.

[0089] FIG. 5 is a diagram showing a comparison example between the generations retained when applying a backup method of deleting from old data and the generations retained when applying the diversity concentration method.

[0090] FIG. 5(1) is a diagram showing the update history by backup. The blocks shown in FIG. 5(1) represent all the generations that have been backed up. Among the blocks shown in FIG. 5(1), the region marked with symbol a indicates the data of the 41st to 50th generations. Among the blocks shown in FIG. 5(1), the region marked with symbol b indicates the data of the 31st to 40th generations. Among the blocks shown in FIG. 5(1), the region marked with symbol c indicates the data of the 21st to 30th generations. Among the blocks shown in FIG. 5(1), the region marked with symbol d indicates the data of the 11th to 20th generations. Among the blocks shown in FIG. 5(1), the region marked with symbol e indicates the data of the 1st to 10th generations.

[0091] FIG. 5(2) is a diagram showing an example of the data retained when applying a backup method of deleting from old data. FIG. 5(2) shows an example of the case where data for up to 10 generations can be retained.

[0092] When applying the backup method of deleting from old data shown in FIG. 5(2), in order to increase the number of generations that can be retained, it is assumed that the user has voluntarily limited the backup frequency to half.

[0093] Of FIG. 5(2), FIG. 5(2-1) is a diagram showing an example of data retained when 20 generations have passed since the start of backup. FIG. 5(2-1) shows an example in which data from approximately the 1st to 20th generations is retained skipping one generation at a time. Of FIG. 5(2), FIG. 5(2-2) is a diagram showing an example of data retained when 50 generations have passed since the start of backup. FIG. 5(2-2) shows an example in which data from approximately the 31st to 40th generations is retained skipping one generation at a time.

[0094] As shown in FIG. 5(2-2), when applying a backup method that deletes old data, even if the backup frequency is limited to half in order to increase the number of generations that can be retained, all data before a certain point (in this example, data from generations earlier than the 30th generation) will be lost.

[0095] FIG. 5(3) is a diagram showing an example of data retained when applying the diversity concentration method using the diversity concentration device 10 of the present embodiment. FIG. 5(3) shows an example where, similar to FIG. 5(2), data for up to 10 generations can be retained.

[0096] When applying the diversity concentration method, the diversity concentration device 10 automatically retains data in a state with high diversity. Therefore, when applying the diversity concentration method, unlike when applying a backup method that deletes old data, there is no need for the user to autonomously limit the backup frequency.

[0097] Of FIG. 5(3), FIG. 5(3-1) is a diagram showing an example of data retained when 20 generations have passed since the start of backup. FIG. 5(3-1) shows an example in which data from approximately the 1st to 20th generations is retained skipping one generation at a time. Of FIG. 5(3), FIG. 5(3-2) is a diagram showing an example of data retained when 50 generations have passed since the start of backup. FIG. 5(3-2) shows an example in which data from approximately the 1st to 50th generations is retained skipping four generations at a time.

[0098] As shown in Fig. 5(3-2), when the diversity concentration method is applied, even if the backup frequency is not autonomously restricted, the data of the first generation at the beginning of backup start can be retained.

[0099] Also, as shown in Fig. 5(3-2), when the diversity concentration method is applied, even at the time when 50 generations have passed since the start of backup, generally the data of the first to 50th generations are retained skipping four generations. Therefore, for example, even if infected by malware with a long latency period, it is certain that the data of generally 0 to 4 generations before just before infection are retained. Thus, when the diversity concentration method is applied, the diversity of the retained generations can be amplified and diversity can be preserved.

[0100] Next, the transition of the amount of information to be retained will be described. Fig. 6 is a diagram showing a comparison example of the transition of the amount of information retained when the backup method of deleting old data is applied and the transition of the amount of information retained when the diversity concentration method is applied.

[0101] Fig. 6(1) is a diagram showing the transition of the amount of information retained when the backup method of deleting old data is applied. Fig. 6(2) is a diagram showing the transition of the amount of information retained when the diversity concentration method using the diversity concentration device 10 of the present embodiment is applied. Fig. 6(3) is a diagram showing the comparison between the example shown in Fig. 6(1) and the example shown in Fig. 6(2). As shown in Fig. 6(3), the total amount of information S1 retained in the example shown in Fig. 6(1) and the total amount of information S2 retained in the example shown in Fig. 6(2) are the same amount of information.

[0102] As shown in Fig. 6(1), when the backup method of deleting old data is applied, all the information before the time T1 has passed since being stored is retained, but no information after the time T1 has passed is retained.

[0103] As shown in Fig. 6(2), when the diversity concentration method is applied, although the amount of information retained gradually decreases with the passage of time, the information after the time T1 has passed is also retained.

[0104] When applying the diversity concentration method, unlike the memory characteristic of completely losing the memory before a certain point in time as shown in Fig. 6(1), it is possible to reproduce a forgetting curve that gently loses memory over time as shown in Fig. 6(2).

[0105] The backup method of deleting from old data has the following characteristics (1) and (2). (1) The retroactive period is limited to the product of the number of generations x c [times] that can be stored in the storage and the backup interval s [h], that is, x c ·s [h]. It is not possible to obtain data before that period. (2) When data before the time t [h] to be retroactively retrieved is required within the retroactive period, data from t to t + s [h] can be obtained. The range of the time difference does not depend on t.

[0106] The diversity concentration method according to the present invention has the following characteristics (1) and (2). (1) It is retroactive to the beginning of operation regardless of the capacity x c of the storage or the backup interval s [h]. (2) When data before the time t [h] to be retroactively retrieved is required, data from approximately t to t·(1 + ε) [h] can be obtained. The range of the time difference depends on t.

[0107] The time difference rate ε described in the characteristic (2) of the diversity concentration method according to the present invention is obtained, for example, as follows.

[0108] Taking the case where the number of generations x c that can be stored in the storage is 100 generations and the backup interval s is 5 hours as an example, and the case of operating for 4096 hours is described. The transition point x t in this case is 21 generations ago. Therefore, the ideal distribution is calculated by the following formula (4). The base in formula (4) corresponds to 1 + ε. Therefore, the time difference rate ε in this case is 4.75%.

[0109] [Number]

[0110] Next, taking as an example the case where the number of generations x that can be stored in the storage is 100 generations and the backup interval s is 5 hours, and the operation is performed for 262144 hours. The transition point x c in this case is 10 generations ago. Therefore, the ideal distribution is calculated by the following formula (5). Thus, the time difference rate ε in this case is 9.98%. t As described above, when the diversity concentration method according to the present invention is applied, even if the operation time and the backup frequency are increased by 64 times, the time difference rate and the like do not increase by 64 times. The influence when the operation time and the backup frequency are increased by 64 times is only a slight increase in the time difference rate. For example, assume that the latency period of the malware is about half a year. When the diversity concentration method according to the present invention is applied, data within the range of the time difference within several percent of the time point when the malware lurks can be obtained. If this time difference rate is acceptable, it can be said that the backup method applying the diversity concentration method according to the present invention, which has no restriction on the retrievable period, is extremely practical.

[0111]

Number

[0112] As described above, when the diversity concentration method according to the present invention is applied, even if the operation time and the backup frequency are increased by 64 times, the time difference rate and the like do not increase by 64 times. The influence when the operation time and the backup frequency are increased by 64 times is only a slight increase in the time difference rate. For example, assume that the latency period of the malware is about half a year. When the diversity concentration method according to the present invention is applied, data within the range of the time difference within several percent of the time point when the malware lurks can be obtained. If this time difference rate is acceptable, it can be said that the backup method applying the diversity concentration method according to the present invention, which has no restriction on the retrievable period, is extremely practical.

[0113] FIG. 7 is a diagram showing a comparison example between the data distribution retained when the backup method of deleting old data is applied and the data distribution retained when the diversity concentration method is applied.

[0114] FIG. 7(1-1) is a diagram showing an example where 4096 generations have passed since the start of backup in a configuration where the backup method of deleting old data is applied. FIG. 7(1-2) is a diagram showing an example where 4096 generations have passed since the start of backup in a configuration where the diversity concentration method is applied.

[0115] FIG. 7(2-1) is a diagram showing an example in which 262,144 generations have elapsed since the start of backup in a configuration to which a backup method of deleting from old data is applied. FIG. 7(2-2) is a diagram showing an example in which 262,144 generations have elapsed since the start of backup in a configuration to which a diversity concentration method is applied.

[0116] As shown in FIG. 7, when the diversity concentration method is applied, data can always be held in a state approximating the ideal distribution as compared with the case where the backup method of deleting from old data is applied. Further, when the time elapses from 4,096 generations to 262,144 generations, the deviation from the ideal distribution becomes larger in the case where the backup method of deleting from old data is applied. However, when the diversity concentration method is applied, there is no significant change in the state approximating the ideal distribution.

[0117] Some or all of the above embodiments may be described as follows in (1) to (2), but are not limited thereto.

[0118] (1) Specific means (for example, realized by exclusion target specifying means 6) for specifying data to be excluded such that the distribution of attribute values in a group of data (for example, a set of data itself, a set of metadata, a set of catalog data, etc.), which is a set of data including attribute values (for example, those indicating the attributes of the data itself, those indicating the attributes of other data, those indicating the attributes of objects or events not represented by data, etc.), approximates a specific distribution (for example, an ideal distribution) after the data to be excluded is excluded from the group of data; and output means (for example, realized by output means 7) for outputting information indicating the data to be excluded specified by the specific means. With such a configuration, data with poor contribution to diversity can be suitably specified as data to be excluded from a group of data, and the diversity of the group of data can be efficiently maintained.

[0119] (2) In the diversity concentration device described in (1) above, the specific means specifies a specific distribution based on a predetermined total amount defined for a group of data (for example, the upper limit total number of a group of data. The oldest point c is specified by the upper limit total number) (for example, the ideal distribution specifying means 61 corresponds to the part that executes the process of step S3 shown in FIG. 4), and specifies a predetermined distribution (for example, the post-change distribution), which is the distribution of attribute values when any one data is excluded from a group of data (for example, the post-change distribution evaluation means 62 corresponds to the part that executes the process of step S6 shown in FIG. 4). When specifying the predetermined distribution with the highest degree of coincidence with the specific distribution or the predetermined distribution with the lowest degree of non-coincidence with the specific distribution, the data excluded is specified as the data to be excluded (for example, the exclusion target specifying means 6 corresponds to the part that executes the processes of steps S10 to S11 shown in FIG. 4). With such a configuration, it is possible to specify the data to be excluded so as to approximate the most desirable distribution state of the attribute values.

[0120] As described above, the present invention has been described with reference to the embodiments, but the present invention is not limited to the above embodiments. Various changes that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention.

Industrial Applicability

[0121] The present invention is applicable to applications for version management of CAD (computer-assisted / aided drafting), programming, game save data, etc. For example, by applying the present invention, in version management, it becomes possible to extend the storage period and diversify the storage time. Also, by making it possible to extend the storage period and diversify the storage time, when an irreversible operation error occurs, it becomes easy to recover to a state before the occurrence of the operation error and with a slight rollback. Therefore, the present invention is also preferably applicable to the generation management of backup data against latent malware.

[0122] The present invention is applicable to applications for managing data recorded by security cameras and event data recorders. By applying the present invention, unlike the conventional usage forms with limited storage periods, the storage period can be extended. Therefore, it is possible to suppress the occurrence of a virtual statute of limitations where investigation becomes difficult due to the absence of recorded data. Also, by applying the present invention, the usage form changes from the conventional one that records the recent situation to one that can record the normal state. These effects are expected to lead to an improvement in the deterrence against crimes and harassment.

[0123] Moreover, by applying the present invention, for recorders that strongly require recording of the normal state such as life logs and animal logs, it becomes possible to preserve the normal state recording without expending the labor of data management.

[0124] The present invention is applicable to applications where artificial intelligence realizes more human-like memory characteristics. For example, unlike the conventional memory characteristics in which memories before a certain period are completely lost, the memory system to which the present invention is applied can reproduce a forgetting curve in which memories are gradually lost over time.

[0125] The present invention is applicable to applications that cause automatic backup of a file and automatic deletion of backup data each time the file is saved, triggered by overwriting and saving the file.

[0126] The present invention is suitably applied to versioning file systems, version management systems, digital cameras, and event data recorders, which are software and information processing devices for accumulating and managing various data. In this case, even if the free capacity of the hard disk that was conventionally wasted is allocated to these data, the free capacity can be secured at any time as needed, so that the memory resources can be utilized to the maximum extent.

[0127] The present invention is applicable to applications that provide a decision-making support function for maximizing diversity in a group allowed by limited resources other than memory resources. For example, by incorporating this method as decision-making support for production and inventory management, etc., it is possible to enrich the variety of limited inventory.

[0128] The present invention is applicable to applications that provide a diversification mechanism that does not necessarily depend on probability for information processing devices and the like. For example, in heuristics for optimizing evaluation values such as genetic algorithms, there is a problem (initial convergence problem) that the solution group loses diversity and it is difficult to obtain a global optimal solution. The mechanism to which the present invention is applied can output the contribution degree of each element of the solution group with respect to the diversity of arbitrary attribute values. This contribution degree can be regarded as an additional evaluation value with respect to the evaluation value inherent to the heuristic. Therefore, it is expected that the mechanism to which the present invention is applied can preserve diversity and reduce the initial convergence problem. Such a mechanism is applicable to applications such as circuit design, material development, and drug development (for example, drug discovery and pharmaceutical manufacturing). That is, when searching for gunpowder or an alloy according to the intended application, it is expected that a global optimal solution can be obtained from a variety of solution groups by treating parameter sets such as optimal arrangement, shape, and blending ratio as a data group. At the same time, if it is applied to in-silico drug discovery for searching for new drugs by simulating molecular structures and the like with a computer, it can be utilized for research on drugs with higher activity and fewer side effects.

Explanation of Signs

[0129] 1 Memory unit 2 Monitoring means 3 Extraction means 4 Exclusion means 5 Input means 6 Exclusion target specifying means 61 Ideal distribution specifying means 62 Changed distribution evaluation means 7 Output means 10 Diversity concentration device 20 User terminal

Claims

1. Specific means for identifying data to be excluded such that the distribution of attribute values in the group of data after the data to be excluded is excluded from a group of data that is a set of data including attribute values approximates a specific distribution; Output means for outputting information indicating the data to be excluded identified by the specific means. A diversity enrichment device characterized by the above.

2. The specific means: Identifies the specific distribution based on a predetermined total amount defined for the group of data; Identifies a predetermined distribution that is the distribution of attribute values when any one data is excluded from the group of data; Identifies, as the data to be excluded, the data excluded when identifying the predetermined distribution having the highest degree of coincidence with the specific distribution or the predetermined distribution having the lowest degree of non - coincidence with the specific distribution. The diversity enrichment device according to Claim 1.

3. A computer: Identifies data to be excluded such that the distribution of attribute values in the group of data after the data to be excluded is excluded from a group of data that is a set of data including attribute values approximates a specific distribution; Outputs information indicating the identified data to be excluded. A diversity enrichment method characterized by the above.

4. A diversity enrichment program for causing a computer to: Execute a specific process for identifying data to be excluded such that the distribution of attribute values in the group of data after the data to be excluded is excluded from a group of data that is a set of data including attribute values approximates a specific distribution; Execute an output process for outputting information indicating the data to be excluded identified by the specific process. ​