Information processing method, information processing device, and information processing program

By associating genomic genotypes with data density and calculating incentives based on rarity and attributes, the method efficiently collects rare genetic data for SNP genotype imputation, improving the accuracy of genome-wide association studies.

JP7723012B2Active Publication Date: 2025-08-13PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022572930
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-28
Filing Date
2021-11-10
Publication Date
2025-08-13
Estimated Expiration
2041-11-10

AI Technical Summary

Technical Problem

Existing technologies fail to efficiently collect rare genetic data necessary for SNP genotype imputation, which is crucial for genome-wide association studies, due to the lack of consideration for data density and user incentives.

Method used

An information processing method that associates genomic genotypes with data density, identifies the rarity of genetic data based on its location, and calculates incentives based on rarity and attribute information to motivate users to provide rare genetic data.

Benefits of technology

This approach enables efficient collection of rare genetic data by providing higher incentives to users who provide data with low density and valuable attributes, enhancing the accuracy and efficiency of SNP genotype imputation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007723012000001
    Figure 0007723012000001
  • Figure 0007723012000002
    Figure 0007723012000002
  • Figure 0007723012000003
    Figure 0007723012000003
Patent Text Reader

Abstract

This information processing device 1 comprises: an acquisition unit (121) that acquires gene data including a base sequence indicating the genotype of a user, said gene data being detected by a gene detection device; a region specifying unit (122) that specifies a region of reference data in which the gene data is positioned; a degree-of-rarity calculation unit (123) that calculates the degree of rarity indicating the rarity of the gene data, on the basis of the data density associated with the specified region; an incentive calculation unit (125) that calculates an incentive to be awarded to a user in accordance with the calculated degree of rarity; and an output unit (126) that outputs the calculated incentive.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to techniques for collecting genetic data. [Background technology]

[0002] In recent years, a technology called SNP genotype imputation has become known that estimates genotypes in regions that cannot be obtained using SNP (single nucleotide polymorphism) microarrays. SNP genotype imputation uses reference data that contains high-density information indicating SNP genotypes. To build high-density reference data, it is necessary to efficiently collect genetic data from regions with low data density, i.e., rare genetic data, rather than collecting genetic data haphazardly.

[0003] Patent Document 1 discloses a method for providing bioinformation data that uses blockchain technology to make it difficult to expose bioinformation data and to falsify or alter genome data.

[0004] Patent document 2 discloses an information trading device that presents a reward amount to an information provider and then provides only user information corresponding to the information provider who has given consent to the information user, and adjusts the reward amount depending on the acquisition status of the user information.

[0005] However, none of the above conventional techniques take into consideration the efficient collection of rare genetic data, and therefore further improvement is needed. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Patent No. 6661742 [Patent Document 2] Patent No. 5978198 Summary of the Invention

[0007] The present disclosure has been made to solve the above-mentioned problems, and aims to provide a technology that can efficiently collect rare genetic data.

[0008] An information processing method according to one aspect of the present disclosure is an information processing method in an information processing device that processes information using reference data, wherein the reference data is data in which a base sequence indicating a genomic genotype is pre-associated with a data density according to the locus of the base sequence, and the reference data is detected by a gene detection device to obtain genetic data including a base sequence indicating a user's genotype, identify an area in the reference data in which the genetic data is located, calculate a rarity indicating the rarity of the genetic data based on the data density associated with the identified area, calculate an incentive to be granted to the user based on the calculated rarity, and output the calculated incentive.

[0009] According to the present disclosure, rare genetic data can be collected efficiently. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a diagram illustrating an example of an overall configuration of an information processing system to which an information processing device according to a first embodiment of the present disclosure is applied. [Figure 2] 2 is a block diagram showing an example of the configuration of the information processing device shown in FIG. 1. FIG. [Figure 3] FIG. 1 is an explanatory diagram of terms related to genetic analysis. [Figure 4] FIG. 2 is a diagram illustrating an example of a data configuration of reference data. [Figure 5] FIG. 1 shows reference data represented according to data density. [Figure 6] 4 is a flowchart showing an example of processing performed by the information processing device according to the first embodiment of the present disclosure. [Figure 7] FIG. 10 is a block diagram showing an example of a configuration of an information processing device according to a second embodiment of the present disclosure. [Figure 8] FIG. 2 is a diagram illustrating an example of a data configuration of area reference data. [Figure 9] FIG. 9 is a diagram showing the area reference data shown in FIG. 8 in accordance with data density. [Figure 10] 10 is a flowchart showing an example of processing by an information processing device according to a second embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0011] (Background to this disclosure) Genome-wide association studies are currently being conducted on hundreds of thousands of people to identify the genotypes of tens of millions of SNPs across the entire human genome and evaluate the association between target traits and SNP genotypes. Genome-wide association studies require tens of millions of SNP genotypes. Meanwhile, SNP microarrays, which allow for low-cost and easy SNP genotyping, have become widespread in recent years.

[0012] Since SNP microarrays can only obtain genotypes for a few hundred thousand SNPs, the genetic data obtained by SNP microarrays cannot be directly applied to genome-wide association studies. Therefore, SNP genotype imputation is used to statistically infer the genotypes of tens of millions of SNPs from the genetic data obtained by SNP microarrays.

[0013] In SNP genotype imputation, the genotypes of SNPs in unobserved regions are inferred by interpolating the base sequences of genetic data obtained by SNP microarray with the base sequences in the reference data. However, SNP genotype imputation requires reference data containing SNP genotypes at a high density. To achieve this, rather than collecting genetic data haphazardly, it is necessary to efficiently collect genetic data corresponding to regions with low data density, i.e., rare genetic data.

[0014] The above-mentioned Patent Document 1 merely discloses that a second user's biometric information data encrypted with the second user's public key is provided to a first user who has successfully passed user authentication using blockchain technology, and its only objective is to prevent the exposure of biometric information data and the forgery or falsification of genome data. Therefore, Patent Document 1 does not allow for efficient collection of rare genetic data.

[0015] In the above-mentioned Patent Document 2, the user information provided by the information provider is personal information including location information, atmospheric pressure information, sound pickup information, illuminance information, frequency information, as well as age, occupation, and annual income, and is not genetic data. Therefore, Patent Document 2 cannot determine an appropriate incentive to be given to the information provider depending on the rarity of the genetic data, and as a result, it is not possible to efficiently collect rare genetic data.

[0016] Therefore, the present inventors have come up with the following aspects of the present disclosure in order to efficiently collect rare genetic data.

[0017] An information processing method according to one aspect of the present disclosure is an information processing method in an information processing device that processes information using reference data, wherein the reference data is data in which a base sequence indicating a genomic genotype is pre-associated with a data density according to the locus of the base sequence, and the reference data is detected by a gene detection device to obtain genetic data including a base sequence indicating a user's genotype, identify an area in the reference data in which the genetic data is located, calculate a rarity indicating the rarity of the genetic data based on the data density associated with the identified area, calculate an incentive to be granted to the user based on the calculated rarity, and output the calculated incentive.

[0018] According to this configuration, the region in the reference data where the genetic data provided by the user is located is identified, and the rarity of the genetic data is calculated based on the data density associated with the identified region. Then, an incentive to be given to the user is calculated according to the rarity, and the calculated incentive is output. Therefore, it is possible to provide a higher incentive to a user who provides rare genetic data than to a user who provides genes with a low rarity. As a result, rare genetic data can be collected efficiently.

[0019] In the above information processing method, the genetic data may be associated with attribute information including user attributes, and further, the contribution of the genetic data to genetic analysis may be calculated based on the attribute information, and the incentive may be calculated based on the rarity and the contribution.

[0020] When genetic analysis is performed using genetic data, the availability of attribute information on the user who provided the genetic data increases the likelihood of obtaining useful genetic analysis results. According to this configuration, the degree of contribution to the genetic analysis is calculated based on the attribute information, and the incentive is calculated by further taking the calculated degree of contribution into consideration. This motivates users to provide attribute information, enabling efficient collection of genetic data associated with useful attribute information.

[0021] In the above information processing method, the genetic data may be associated with locus information indicating the locus of a base sequence indicating the genotype, and the rarity may be calculated by identifying the region in the reference data in which the genetic data is located based on the locus information.

[0022] According to this configuration, the genetic data is associated with locus information indicating the locus of the gene, so that the region in the reference data where the genetic data is located can be easily identified.

[0023] In the above information processing method, the attribute information may include information indicating the user's place of residence, the reference data may include a plurality of regional reference data corresponding to predetermined regions, and the identification of the region may include identifying the region in which the genetic data is located in the regional reference data corresponding to the information regarding the place of residence.

[0024] Since the genotypes of users living in the same region tend to be similar, performing SNP genotype imputation using regional reference data according to the region can improve the estimation accuracy. In this case, the genetic data of users who reside in regions corresponding to regional reference data with low data density is rarer than the genetic data of users who reside in regions corresponding to regional reference data with high data density. This configuration makes it possible to calculate incentives according to the region of residence of the user who provided the genetic data. Therefore, genetic data that is rare from a regional perspective can be efficiently collected.

[0025] In the above information processing method, the degree of contribution may be calculated by determining whether the attribute information includes information indicating the user's blood relationship, and if it is determined that the attribute information includes information indicating the blood relationship, the degree of contribution may be calculated to be higher than if it is determined that the information indicating the blood relationship does not include.

[0026] According to this configuration, if the attribute information includes information indicating the blood relationship of the user, it is possible to provide a higher incentive to the user, thereby motivating the user to provide information indicating the blood relationship that is useful in genetic analysis, and enabling efficient collection of information indicating the blood relationship.

[0027] In the information processing method, in calculating the degree of contribution, the degree of contribution may be calculated to be higher as the amount of information indicating the blood relationship included in the attribute information increases.

[0028] According to this configuration, it is possible to provide a user with a higher incentive as the amount of information indicating blood relationships increases, thereby enabling efficient collection of substantial information indicating blood relationships.

[0029] In the above information processing method, the degree of contribution may be calculated by determining whether the attribute information includes information indicating the user's lifestyle pattern, and if it is determined that the attribute information includes information indicating the lifestyle pattern, the degree of contribution may be calculated to be higher than if it is determined that the attribute information does not include information indicating the lifestyle pattern.

[0030] According to this configuration, if the attribute information includes a user's lifestyle patterns, it is possible to provide a higher incentive to the user, thereby motivating the user to provide lifestyle pattern data that will be useful in epigenetics research, and enabling efficient collection of lifestyle pattern data.

[0031] In the information processing method, in calculating the degree of contribution, the degree of contribution may be calculated to be higher as the amount of information indicating the lifestyle pattern of the user included in the attribute information increases.

[0032] According to this configuration, it is possible to provide a higher incentive to the user as the amount of information indicating a lifestyle pattern increases, thereby enabling efficient collection of information indicating a lifestyle pattern with substantial content.

[0033] An information processing device according to another aspect of the present disclosure is an information processing device that processes information using reference data, wherein the reference data is data in which a base sequence indicating a genotype of a genome is pre-associated with a data density according to the locus of the base sequence, and the information processing device is equipped with an acquisition unit that acquires genetic data detected by a gene detection device and including a base sequence indicating a user's genotype, a region identification unit that identifies a region in the reference data in which the genetic data is located, a rarity calculation unit that calculates a rarity indicating the rarity of the genetic data based on the data density associated with the region identified by the region identification unit, an incentive calculation unit that calculates an incentive to be granted to the user based on the rarity calculated by the rarity calculation unit, and an output unit that outputs the incentive calculated by the incentive calculation unit.

[0034] An information processing program according to yet another aspect of the present disclosure is an information processing program that causes a computer to function as an information processing device that processes information using reference data, wherein the reference data is data in which a base sequence indicating a genotype of a genome is pre-associated with a data density according to the locus of the base sequence, and causes the computer to function as an acquisition unit that acquires genetic data detected by a gene detection device and including a base sequence indicating a user's genotype, an area identification unit that identifies an area in the reference data in which the genetic data is located, a rarity calculation unit that calculates a rarity indicating the rarity of the genetic data based on the data density associated with the area identified by the area identification unit, an incentive calculation unit that calculates an incentive to be granted to the user based on the rarity calculated by the rarity calculation unit, and an output unit that outputs the incentive calculated by the incentive calculation unit.

[0035] The present disclosure can also be realized as an information processing system operated by such an information processing program. Needless to say, the information processing program can be distributed on a computer-readable non-transitory recording medium such as a CD-ROM or via a communication network such as the Internet.

[0036] Note that each of the embodiments described below represents a specific example of the present disclosure. The numerical values, shapes, components, steps, and step orders shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concept are described as optional components. Furthermore, in all of the embodiments, the respective contents can be combined.

[0037] (Embodiment 1) 1 is a diagram showing an example of the overall configuration of an information processing system to which an information processing device 1 according to the first embodiment of the present disclosure is applied. The information processing system includes an information processing device 1, a providing terminal 2, and a user terminal 3. The information processing device 1 to the user terminal 3 are connected to each other via a network NT so as to be able to communicate with each other.

[0038] The information processing device 1 is configured, for example, as a cloud server including one or more computers. The information processing device 1 receives genetic data provided by a user from a provision terminal 2, and calculates an incentive to be given to the user based on the received genetic data.

[0039] The provision terminal 2 is configured, for example, by a computer owned by a medical institution, and transmits genetic data to the information processing device 1. The genetic data is detected by a gene detection device and contains a base sequence that indicates the user's genotype. An SNP microarray, for example, can be used as the gene detection device. In an SNP microarray, DNA fragments called probes that detect base differences are densely packed on a chip. The SNP microarray detects the genotypes of hundreds of thousands of SNPs. The gene detection device is not limited to an SNP microarray, and other devices may also be used.

[0040] Genetic data is associated with a user identifier that identifies the user who provided the genetic data. Furthermore, genetic data is associated with locus information that indicates the locus of the base sequence that indicates the SNP genotype. This locus information indicates the locus on the genome of the base sequence that indicates the SNP genotype.

[0041] The user terminal 3 is an information processing device owned by a user who provides genetic data. In detail, the user terminal 3 is configured as, for example, a mobile information terminal such as a smartphone or a tablet terminal, or a stationary computer such as a laptop computer. The user terminal 3 acquires attribute information input by the user and transmits the acquired attribute information to the information processing device 1.

[0042] The network NT is composed of a wide area communication network including, for example, the Internet and a mobile phone communication network.

[0043] Here, the genetic data is transmitted from the providing terminal 2 to the information processing device 1, but the present disclosure is not limited to this, and the genetic data may be transmitted from the user terminal 3 to the information processing device 1. In this case, the user terminal 3 may acquire the genetic data detected by the SNP microarray, associate it with the attribute information, and transmit it to the information processing device 1. Alternatively, the attribute information may be transmitted from the providing terminal 2. In this case, the providing terminal 2 may acquire the genetic data detected by the SNP microarray, associate it with the attribute information, and transmit it to the information processing device 1.

[0044] 2 is a block diagram showing an example of the configuration of the information processing device 1 shown in FIG. 1. The information processing device 1 includes a communication unit 11, a processor 12, and a memory 13. The communication unit 11 is configured with a communication circuit for connecting the information processing device 1 to a network NT. The communication unit 11 receives genetic data transmitted from the provision terminal 2. The genetic data received here is associated with a user identifier and locus information. The communication unit 11 receives attribute information transmitted from the user terminal 3. The received attribute information is associated with a user identifier.

[0045] The memory 13 is configured by a non-volatile storage device such as an SSD (Solid State Drive) or an HDD (Hard Disc Drive). The memory 13 stores reference data 131 and incentive information 132.

[0046] The reference data 131 is reference data used in genotype imputation, and is data in which a base sequence indicating the genotype of a human genome is associated with a data density according to the locus of the base sequence.

[0047] Here, we will explain the terminology used in genetic analysis. Figure 3 is an explanatory diagram of the terminology related to genetic analysis. In Figure 3, two lines indicate homologous chromosomes 401 and 402. Locus 403 indicates the location of a gene on homologous chromosomes 401 and 402. Allele 404 refers to a pair of genes on homologous chromosomes 401 and 402. Genotype 405 refers to a combination of alleles 404. Haplotype 406 refers to a combination of alleles 404. Diplotype 407 refers to a combination of haplotypes 406.

[0048] Next, a specific example of the reference data 131 will be described. Fig. 4 is a diagram showing an example of the data configuration of the reference data 131. In the example of Fig. 4, the reference data 131 has a data structure in which two base sequences corresponding to homologous chromosomes 401 and 402 are arranged in a meandering pattern in units of two lines. For example, the base sequences are arranged such that the base sequence of the homologous chromosome 401 is arranged in the first line, the base sequence of the homologous chromosome 402 is arranged in the second line, the base sequence continuing from the first line is arranged in the third line, and the base sequence continuing from the second line is arranged in the fourth line.

[0049] Furthermore, in the reference data 131, each locus 403 of the base sequence is associated with a data density. The data density is a value determined according to the number of pieces of data used to determine the base at a certain locus 403. For example, if the number of pieces of data used is 10,000, the data density is set to "1.0," and if the number of pieces of data is 3,000, the data density is set to "0.3." In this way, the reference data 131 is configured as a set of the base sequence of the homologous chromosome 401 and the base sequence of the homologous chromosome 402. Therefore, the reference data 131 includes information indicating genotypes such as alleles, haplotypes, and diplotypes. Note that the reference data 131 may represent tens of millions of base sequences of genes in the human genome, the base sequence of the entire human genome, or the base sequences of tens of millions of SNPs.

[0050] FIG. 5 is a diagram showing the reference data 131 according to data density. In the example of FIG. 5, loci with higher data density are displayed with higher density. For example, the genotype included in the high-density region indicated by reference numeral 601 has its base sequence determined using more data than the genotype included in the low-density region indicated by reference numeral 602. As such, it can be seen that the data density of the reference data 131 varies depending on the locus.

[0051] Next, SNP genotype imputation will be explained. Genetic data detected by SNP microarrays is data in which part of the base sequence of one homologous chromosome and part of the base sequence of the other homologous chromosome are confirmed, with the remaining parts missing, such as "...A...A...A..." and "...G...C...A...." The "..." parts indicate unconfirmed base sequences, with A representing adenine, G representing guanine, and C representing cytosine. SNP genotype imputation predicts the genotype of the SNP in this missing part using reference data 131.

[0052] In SNP genotype imputation, a base sequence pattern determined in genetic data is compared with a base sequence pattern in reference data 131, and a region in reference data 131 where both patterns best match is searched for. Then, the base sequence of a missing part in the genetic data is inferred from the base sequence of reference data 131 in the searched region, and the genotype of the SNP is inferred based on the inference result. The inferred genotype result obtained here is expressed as a probability, for example, for a certain SNP, such as 0.95 for "AA" type, 0.44 for "AG" type, and 0.01 for "GG" type.

[0053] See Fig. 2. The incentive information 132 is information in which, for each of one or more users, a user identifier is associated with an incentive granted to the user. The incentive may be data having economic value, such as electronic money, mileage points, virtual currency, points for purchasing products, or coupons, or may be data having no economic value, such as a certificate.

[0054] The processor 12 is configured by, for example, a CPU, and includes an acquisition unit 121, an area identification unit 122, a rarity calculation unit 123, a contribution calculation unit 124, an incentive calculation unit 125, and an output unit 126. These blocks included in the processor 12 are realized by the CPU executing an information processing program.

[0055] The acquisition unit 121 acquires the genetic data transmitted from the provider terminal 2 using the communication unit 11. The acquisition unit 121 receives the attribute information transmitted from the user terminal 3 using the communication unit 11. The acquisition unit 121 associates the genetic data with the attribute information using the user identifier as a key. This allows for the acquisition of a data set in which the user identifier, genetic data, locus information, and attribute information are associated with each other.

[0056] The attribute information includes personal information of the user, residence information indicating the user's residence, blood relationship information indicating the user's blood relationship, and lifestyle pattern information indicating the user's lifestyle pattern.

[0057] The user's personal information includes the user's age, gender, occupation, etc. The user's personal information is information obtained, for example, by the user inputting it into the user terminal 3. The residence information includes information indicating the name of the area where the user resides. Here, the name of the area where the user resides includes, for example, at least one of a country name, a prefecture name, and a state name. Note that the information indicating the name of the area where the user resides may include information with greater granularity than a prefecture (for example, Honshu, Shikoku, Kyushu, and Hokkaido in Japan), or may include information with greater granularity than a country (for example, Asia continent, Africa continent, North America continent). The residence information may be obtained by the user inputting it into the user terminal 3, or may be determined based on location data detected by a GPS sensor included in the user terminal 3.

[0058] The lifestyle pattern information indicates, for example, the lifestyle pattern of a user over a predetermined period (e.g., one day). The lifestyle pattern information includes, for example, the average number of cigarettes smoked per day, the average amount of alcohol consumed per day, the average calories burned per day, the average calories consumed per day, the number of meals eaten per day, meal times, the average time of waking up, the average time of going to bed, and the average amount of sleep per day. The lifestyle pattern information may be information input by the user or may be information monitored by a biosensor such as a smartwatch.

[0059] The region identifying unit 122 identifies a region in the reference data where the genetic data acquired by the acquiring unit 121 is located. Here, the region identifying unit 122 may identify the region where the genetic data is located based on locus information associated with the genetic data.

[0060] The rarity calculation unit 123 calculates a rarity indicating the rarity of genetic data based on the data density associated with the region identified by the region identification unit 122. For example, the rarity calculation unit 123 may calculate an average value of density data from density data associated with all loci in the region identified by the region identification unit 122, and calculate the reciprocal of the calculated average value as the rarity. Alternatively, the rarity calculation unit 123 may calculate an average value of density data associated with the loci of confirmed bases in the region identified by the region identification unit 122, and calculate the reciprocal of the calculated average value as the rarity. This makes it possible to calculate the rarity such that the rarity value increases as the average value of data density in the identified region decreases.

[0061] The contribution calculation unit 124 calculates the contribution of the genetic data to genetic analysis based on the attribute information associated with the genetic data. For example, the contribution calculation unit 124 determines whether the attribute information includes blood relationship information, and if it determines that the blood relationship information is included, calculates a higher contribution than if it determines that the blood relationship information is not included. As the blood relationship information, for example, information identifying blood relatives of the user providing the genetic data can be used. As blood relatives, for example, father, mother, brothers, sisters, grandfather, and other relatives can be used. As information identifying blood relatives, for example, an identifier of the blood relative can be used.

[0062] In this case, the contribution calculation unit 124 may calculate a higher value of the contribution as the amount of information in the blood relationship information increases. For example, the contribution calculation unit 124 may calculate a higher value of the contribution as the number of blood relatives indicated by the blood relationship information included in the attribute information increases.

[0063] In genetic analysis, useful analysis results can be obtained by comparing the genotype of a user with the genotypes of the user's blood relatives. Therefore, in this embodiment, the greater the amount of blood relationship information, the higher the user's contribution is calculated to be.

[0064] Furthermore, contribution degree calculation unit 124 may determine whether or not the attribute information includes the user's lifestyle pattern, and if it is determined that the attribute information includes the user's lifestyle pattern, may calculate the contribution degree to be higher than if it is determined that the lifestyle pattern information does not include the user's lifestyle pattern information. In this case, contribution degree calculation unit 124 may calculate the contribution degree to be higher as the amount of information in the lifestyle pattern information increases. For example, contribution degree calculation unit 124 may determine that the amount of information in the lifestyle pattern information is large as the number of types of data included in the lifestyle pattern information, such as the number of cigarettes smoked per day and the amount of alcohol consumed per day, increases.

[0065] Alternatively, the contribution calculation unit 124 may calculate the final contribution as the sum of the contribution calculated based on the blood relationship information and the contribution calculated based on the lifestyle pattern information. For example, if the final calculated contribution is B, the contribution assigned when blood relationship information is included is B1, and the contribution assigned when a lifestyle pattern is indicated is B2, the contribution calculation unit 124 may calculate the contribution as B=B1+B2. In this case, the value of B1 increases as the amount of information indicating blood relationship increases, and the value of B2 increases as the amount of information regarding lifestyle pattern information increases.

[0066] The incentive calculation unit 125 calculates an incentive to be given to a user so that the value increases as the rarity and contribution increase. For example, if the rarity is A and the contribution is B, the incentive calculation unit 125 may calculate the incentive using the following formula:

[0067] Incentive = α·A + β·B (1) Here, α is a weighting coefficient for rarity, and β is a weighting coefficient for contribution. When emphasis is placed on rarity, coefficient α is set to a value greater than coefficient β, and when emphasis is placed on contribution, coefficient β is set to a value greater than coefficient α.

[0068] The output unit 126 outputs the incentive calculated by the incentive calculation unit 125. Here, the output unit 126 may grant the incentive by registering the calculated incentive in the incentive information 132 of the corresponding user. Furthermore, the output unit 126 may transmit presentation information for presenting the calculated incentive to the user to the user terminal 3 using the communication unit 11.

[0069] Next, a description will be given of processing by the information processing device 1 according to the first embodiment of the present disclosure. Fig. 6 is a flowchart showing an example of processing by the information processing device 1 according to the first embodiment of the present disclosure.

[0070] In step S1, the acquisition unit 121 acquires the genetic data transmitted from the provision terminal 2 using the communication unit 11.

[0071] In step S2, the region identifying unit 122 identifies a region in which the genetic data is located in the reference data 131 based on the locus information associated with the genetic data. In the example of FIG. 4, a region 131a surrounded by a rectangle is identified from the reference data 131.

[0072] In step S3, the rarity calculation unit 123 calculates the average value of the data density in the region identified in step S2, and calculates the reciprocal of the calculated average value as the rarity of the genetic data. In the example of Fig. 4, since the average value of the data density in region 131a is 1.3, 1 / 1.3 is calculated as the rarity.

[0073] In step S4, the contribution calculation unit 124 calculates the contribution based on the attribute information associated with the genetic data. In this case, the contribution calculation unit 124 may increase the value of the contribution as the amount of information indicating blood relationship in the attribute information increases, and may increase the value of the contribution as the amount of information of lifestyle pattern information increases.

[0074] In step S5, the incentive calculation unit 125 inputs the rarity calculated in step S3 and the contribution calculated in step S4 into formula (1) to calculate an incentive according to the rarity and the contribution.

[0075] In step S6, the output unit 126 registers the incentive calculated in step S5 in the incentive information 132 of the user who provided the genetic data, thereby providing the incentive to the user.

[0076] As described above, the information processing device 1 according to the present embodiment can provide a high incentive to a user who provides genetic data that is rare and highly contributing to genetic analysis. As a result, rare genetic data that contributes to genetic analysis can be efficiently collected.

[0077] (Embodiment 2) In the second embodiment, an incentive is calculated taking into consideration the user's place of residence. Fig. 7 is a block diagram showing an example of the configuration of an information processing device 1A in the second embodiment of the present disclosure. In the present embodiment, the same components as those in the first embodiment are assigned the same reference numerals, and description thereof will be omitted.

[0078] In the processor 12A, the area identifying unit 122A identifies the area reference data 1310 corresponding to the area of residence of the user who provided the genetic data, based on the area information included in the attribute information. Then, the area identifying unit 122A identifies the area in which the genetic data is located in the identified area reference data 1310. Note that the details of the process of identifying this area are the same as those in the first embodiment, and therefore will not be described again.

[0079] The memory 13A stores three area reference data 1310 corresponding to areas A, B, and C. In this case, the area identification unit 122A determines to which of areas A to C the residence indicated by the residence information belongs, and identifies the area reference data 1310 corresponding to the area to which it belongs. Here, the memory 13 stores three area reference data 1310, but this is just an example, and the memory 13 may store two area reference data 1310, or four or more area reference data 1310.

[0080] 8 is a diagram showing an example of the data configuration of the area reference data 1310. The area reference data 1310 corresponding to area A is generated based on the genetic data of residents of area A, the area reference data 1310 corresponding to area B is generated based on the genetic data of residents of area B, and the area reference data 1310 corresponding to area C is generated based on the genetic data of residents of area C. The only difference between the area reference data 1310 is the population used to generate them, and the detailed data configuration is the same as that of the reference data 131. In other words, the area reference data 1310 is data in which a base sequence indicating a genotype is associated with a data density according to the locus of the base sequence.

[0081] The granularity of regions A to C may be on a country-by-country basis, or on a regional basis that makes up a country (for example, in Japan, prefectures, or Honshu, Shikoku, Kyushu, and Hokkaido), or on a unit larger than a country (for example, the Asian continent, the African continent, or the North American continent).

[0082] Fig. 9 is a diagram showing the area reference data 1310 shown in Fig. 8 according to data density. As shown in Fig. 9, it can be seen that the data density of the area reference data 1310 differs depending on areas A to C.

[0083] When the genotypes of several thousand people in the Japanese population were examined, clear differences in genotypes were confirmed between the Hokkaido and Honshu regions and the Kyushu and Ryukyu regions. This revealed that the genetic backgrounds of the Japanese population differ between the Hokkaido and Honshu regions and the Kyushu and Ryukyu regions. Therefore, when SNP genotype imputation is performed using the regional reference data 1310 corresponding to the user's place of residence, the accuracy of estimating the user's genotype is improved. Therefore, in the second embodiment, in order to efficiently collect rare genetic data in each of the multiple regional reference data 1310, a high incentive is provided to users residing in rare regions.

[0084] Next, a description will be given of processing of the information processing device 1A according to the second embodiment of the present disclosure. Fig. 10 is a flowchart showing an example of processing of the information processing device 1A according to the second embodiment of the present disclosure. In the flowchart of Fig. 10, the same processes as those in Fig. 6 are denoted by the same reference numerals, and descriptions thereof will be omitted.

[0085] In step S101 following step S1, the area specifying unit 122A specifies the residential area of the user who provided the genetic data from the area information included in the attribute information associated with the genetic data acquired in step S1.

[0086] In step S102, the area identification unit 122A identifies the area reference data 1310 corresponding to the residence identified in step S101. Thereafter, a process is executed in which the identified area reference data 1310 and the genetic data acquired in step S1 are used to calculate and output an incentive to be given to the user.

[0087] 8, if the user's place of residence belongs to region A, region reference data 1310 corresponding to region A is identified, and region 1310a in which the genetic data is located is identified in the identified region reference data 1310. Here, since the average value of data density in region 1310a is 1.3, the rarity is calculated as 1 / 1.3.

[0088] 8, if the user's place of residence belongs to region B, region 1310a in which the genetic data is located is identified in region reference data 1310 for region B. Here, since the average data density in region 1310a is 0.3, the rarity is calculated as 1 / 0.3.

[0089] In the example of FIG. 8, the average data density of area 1310a is greatest in region A, followed by region C and region B. Therefore, the order of rarity is region B, region C and region A. As a result, the incentive given to users belonging to region B is greatest, and the incentive given to users belonging to region A is least.

[0090] In this way, the information processing device 1A in the second embodiment can provide a high incentive to users who reside in areas corresponding to the low data density area reference data 1310. Therefore, it is possible to motivate users who reside in areas corresponding to the low data density area reference data 1310 to provide genetic data, and genetic data can be collected efficiently.

[0091] The present disclosure can employ the following modifications.

[0092] (1) While the region identifying unit 122 identified the region 131a using locus information associated with the genetic data, the present disclosure is not limited to this. For example, the region identifying unit 122 may compare the base sequence pattern of the genetic data with the base sequence pattern of the reference data 131, search for a region in the reference data 131 where the two patterns most closely match, and identify the searched region as the region 131a where the genetic data is located. The same applies to the region identifying unit 122A.

[0093] (2) Although the incentive information 132 is stored in the information processing device 1, the present disclosure is not limited to this. For example, the incentive information 132 may be stored in an external server owned by an administrator who manages the incentives. If the incentive is electronic money, the administrator may be, for example, a financial institution; if the incentive is mileage points, the administrator may be, for example, an airline; and if the incentive is points for purchasing products, the administrator may be, for example, a points management company.

[0094] (3) In the first embodiment, the incentive calculation unit 125 may calculate the incentive based only on the rarity. In this case, the contribution calculation unit 124 is not necessary.

[0095] (4) Although the reference data 131 is stored in the information processing device 1, the present disclosure is not limited to this, and the reference data 131 may be stored in an external server. [Industrial Applicability]

[0096] According to the present disclosure, rare genetic data can be efficiently collected, which is useful in the genetic industry.

Claims

1. An information processing method executed by an information processing device that processes information using reference data, the reference data is data in which a base sequence indicating a genotype of a genome is previously associated with a data density according to a locus of the base sequence, The information processing device, obtaining genetic data including a base sequence that is detected by a genetic detection device and indicates the user's genotype; Identifying a region in which the genetic data is located in the reference data; calculating a rarity indicating the rarity of the genetic data such that the value increases as the average value of the data density associated with the identified region decreases; calculating an incentive to be given to the user so that the value increases as the calculated rarity increases; outputting the calculated incentive; Information processing methods.

2. The genetic data is associated with attribute information including attributes of the user; Furthermore, a contribution of the genetic data to genetic analysis is calculated based on the attribute information; In calculating the incentive, the incentive is calculated so that the value increases as the rarity and the contribution increase. The information processing method according to claim 1.

3. the genetic data is associated with locus information indicating the locus of the base sequence indicating the genotype; In calculating the rarity, a region in which the genetic data is located in the reference data is identified based on the locus information.

3. The information processing method according to claim 1 or 2.

4. the attribute information includes information indicating a place of residence of the user; the reference data includes a plurality of area reference data corresponding to predetermined areas; The region identification identifies a region in which the genetic data is located in regional reference data corresponding to the information about the place of residence.

3. The information processing method according to claim 2.

5. In calculating the degree of contribution, it is determined whether or not the attribute information includes information indicating a blood relationship between the user and the user, and if it is determined that the attribute information includes the information indicating the blood relationship, the degree of contribution is calculated to be higher than if it is determined that the information indicating the blood relationship is not included.

3. The information processing method according to claim 2.

6. In calculating the degree of contribution, the degree of contribution is calculated to be higher as the amount of information indicating the blood relationship included in the attribute information increases.

6. The information processing method according to claim 5.

7. In calculating the degree of contribution, it is determined whether or not the attribute information includes information indicating a lifestyle pattern of the user, and when it is determined that the attribute information includes the information indicating the lifestyle pattern, the degree of contribution is calculated to be higher than when it is determined that the information indicating the lifestyle pattern is not included.

3. The information processing method according to claim 2.

8. In calculating the degree of contribution, the degree of contribution is calculated to be higher as the amount of information indicating the lifestyle pattern of the user included in the attribute information increases.

8. The information processing method according to claim 7.

9. An information processing device that processes information using reference data, the reference data is data in which a base sequence indicating a genotype of a genome is previously associated with a data density according to a locus of the base sequence, an acquisition unit that acquires genetic data detected by the gene detection device and including a base sequence that indicates the user's genotype; a region specifying unit that specifies a region in which the genetic data is located in the reference data; a rarity calculation unit that calculates a rarity indicating the rarity of the genetic data so that the value increases as the average value of the data density associated with the region specified by the region specifying unit decreases; an incentive calculation unit that calculates an incentive to be given to the user so that the value increases as the rarity calculated by the rarity calculation unit increases; an output unit that outputs the incentive calculated by the incentive calculation unit, Information processing device.

10. An information processing program that causes a computer to function as an information processing device that processes information using reference data, the reference data is data in which a base sequence indicating a genotype of a genome is previously associated with a data density according to a locus of the base sequence, an acquisition unit that acquires genetic data detected by the gene detection device and including a base sequence that indicates the user's genotype; a region specifying unit that specifies a region in which the genetic data is located in the reference data; a rarity calculation unit that calculates a rarity indicating the rarity of the genetic data so that the value increases as the average value of the data density associated with the region specified by the region specifying unit decreases; an incentive calculation unit that calculates an incentive to be given to the user so that the value increases as the rarity calculated by the rarity calculation unit increases; causing a computer to function as an output unit that outputs the incentive calculated by the incentive calculation unit; Information processing program.

Citation Information

Patent Citations

  • Dithiolphosphoric acid ester, its preparation, and soil vermin controlling agent containing said ester as active component

    JP1984078198A

  • System and method for using whole-genome information

    JP2020149188A

  • Research support system, research support apparatus, research support method and research support program

    JP2020177566A

  • Method for providing bioinformation data based on multiple blockchains, method for storing bioinformation data, and bioinformation data transmission system

    JP6661742B2

  • Methods and Systems for Genomic Analysis Using Ancestral Data

    US20090099789A1