Information processing device and information processing method

The information processing device addresses information loss and relationship disruption in anonymization by generating a learning model to minimize feature similarity differences, ensuring effective data analysis post-anonymization.

JP7719940B1Active Publication Date: 2025-08-06KDDI CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024215326
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-08-06
Estimated Expiration
2044-12-10

AI Technical Summary

Technical Problem

Anonymization techniques for high-dimensional or complex structured data result in significant information loss and disrupt the relationship between original data and related data, impacting data analysis.

Method used

An information processing device that acquires user data associating user identification information with simpler second data, generates a learning model to output features indicating first data characteristics, and trains the model to minimize similarity differences between features from same and different clusters, allowing for anonymization while preserving data relationships.

Benefits of technology

Preserves data relationships during anonymization, maintaining data integrity for effective analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007719940000001_ABST
    Figure 0007719940000001_ABST
Patent Text Reader

Abstract

An information processing device and an information processing method are provided that process original data while taking into consideration the relationship between the original data to be anonymized and other data. [Solution] The information processing device 1 has an acquisition unit 131 that acquires user data that associates user identification information for each of a plurality of users with first data and second data that is related to the first data and has a simpler data structure than the first data, and a generation unit 132 that inputs two pieces of first data corresponding to two pieces of second data that belong to the same cluster when the second data is clustered into a learning model that outputs features that indicate the characteristics of the first data in response to input of the first data, inputs first data corresponding to two pieces of second data that belong to different clusters into the learning model, and generates a learning model that has been trained so that the similarity between the two features output from the learning model is smaller than the similarity between the two features output from the learning model.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device and an information processing method. [Background technology]

[0002] Conventionally, user information, which is information about users, is collected and analyzed. In this case, in order to protect the privacy of users, at least a part of the user information is anonymized by performing a process such as K-anonymization (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2021-117679 Summary of the Invention [Problem to be solved by the invention]

[0004] Many anonymization techniques, such as K-anonymization, anonymize data by processing the original data. However, if the original data to be processed is high-dimensional or has a complex structure, applying anonymization techniques will result in a significant loss of information from the original data.

[0005] In order to avoid a significant loss of information in the original data, it is possible to convert the original data so as not to lose the meaning of the original data. However, if the original data is converted without considering the relationship between the original data and other data related to the original data, the relationship between the original data and other data will be lost, which may have a negative impact on data analysis using the converted data and other data.

[0006] Therefore, the present invention has been made in consideration of these points, and aims to make it possible to process original data while taking into account the relationship between the original data to be anonymized and other data. [Means for solving the problem]

[0007] An information processing device according to a first aspect of the present invention includes: an acquisition unit that acquires user data that associates user identification information for identifying each of a plurality of users collected by a predetermined business operator, first data, and second data related to the first data, the second data having a simpler data structure than the first data; and a learning model that outputs features indicating the characteristics of the first data in response to input of the first data, the learning model inputting two pieces of first data corresponding to two pieces of second data that belong to the same cluster when the second data is clustered into the learning model, inputting two pieces of first data corresponding to two pieces of second data that belong to different clusters into the learning model, and generating a learning model trained so that the similarity between the two features output from the learning model is smaller than the similarity indicating the degree of similarity between the two features output from the learning model.

[0008] The information processing device may further include a replacement unit that inputs each of a plurality of first data included in the user data into the learning model, acquires features corresponding to each of the plurality of first data, and replaces each of the plurality of first data included in the user data with the acquired features.

[0009] The generation unit may perform K-anonymization on the second data, input into the learning model two pieces of first data that correspond to two pieces of second data that belong to the same cluster when the second data after K-anonymization is clustered, input into the learning model two pieces of first data that correspond to two pieces of second data that belong to different clusters, and generate the learning model trained so that the similarity between the two features output from the learning model is smaller than the similarity between the two features output from the learning model.

[0010] The generation unit may extract the user identification information and the second data as partial data from the user data acquired by the acquisition unit, cluster the partial data, and generate the learning model that has been trained so that the similarity between two feature amounts output from the learning model by inputting two pieces of first data corresponding to two pieces of user identification information belonging to different clusters into the learning model is smaller than the similarity between two feature amounts output from the learning model by inputting two pieces of first data corresponding to two pieces of user identification information belonging to the same cluster into the learning model.

[0011] The first data may be location history data indicating a location history of the user, and the second data may be data related to the location history that is created based on the location history data.

[0012] An information processing method according to a second aspect of the present invention includes the steps of: acquiring user data executed by a computer, the user data associating user identification information for identifying each of a plurality of users collected by a predetermined business operator, first data, and second data related to the first data, the second data having a simpler data structure than the first data; and generating a learning model that outputs features indicating the characteristics of the first data in response to input of the first data, the learning model inputting two pieces of first data corresponding to two pieces of second data that belong to the same cluster when the second data is clustered, inputting two pieces of first data corresponding to two pieces of second data that belong to different clusters into the learning model, and generating a learning model trained so that the similarity between the two features output from the learning model is smaller than the similarity indicating the degree of similarity between the two features output from the learning model. [Effects of the Invention]

[0013] According to the present invention, it is possible to perform processing of original data while taking into consideration the relationship between the original data to be anonymized and other data. [Brief explanation of the drawings]

[0014] [Figure 1] FIG. 1 is a diagram illustrating an overview of an information processing device. [Figure 2] FIG. 2 is a diagram illustrating a functional configuration of an information processing device. [Figure 3] FIG. 3 is a diagram showing an example of a first data group and a second data group. [Figure 4] 10 is a flowchart showing a processing flow in the information processing device. [Figure 5] FIG. 10 is a diagram showing an example in which first data included in user data is converted into feature amounts. [Figure 6] 10 is a flowchart showing a processing flow in the information processing device. DETAILED DESCRIPTION OF THE INVENTION

[0015] [Overview of information processing device 1] 1 is a diagram illustrating an overview of an information processing device 1. The information processing device 1 is a computer that aggregates user data collected by a plurality of businesses. The information processing device 1 is operated by an aggregation business that aggregates data and provides a service of providing the aggregated data, for example, and is communicably connected to external devices such as a business terminal 2 via a communication network (not shown) such as the Internet or a mobile phone line.

[0016] The provider terminal 2 is, for example, a computer operated by a predetermined provider. The provider terminal 2 manages user data that associates a user ID (Identification), which is user identification information for identifying each of a plurality of users collected by the predetermined provider, with first data and second data related to the first data. The second data has, for example, a simpler data structure than the first data.

[0017] The information processing device 1 acquires user data from the operator terminal 2 ((1) in FIG. 1). Although only one operator terminal is shown in FIG. 1, in reality, there are multiple operator terminals corresponding to multiple operators, and the information processing device 1 collects data from each of the multiple operator terminals.

[0018] The information processing device 1 generates a learning model that outputs, in response to input of first data, feature quantities that indicate the characteristics of the first data ((2) in FIG. 1). For example, the information processing device 1 inputs, into the learning model, two pieces of first data that correspond to two pieces of second data that belong to the same cluster when the second data is clustered, and generates a learning model trained so that the similarity between the two feature quantities output from the learning model is smaller than the similarity indicating the degree of similarity between the two feature quantities output from the learning model, by inputting, into the learning model, two pieces of first data that correspond to two pieces of second data that belong to different clusters.

[0019] The information processing device 1 inputs each of the plurality of first data included in the user data into the generated learning model, and acquires feature quantities corresponding to each of the plurality of first data from the learning model ((3) in FIG. 1). Then, the information processing device 1 converts each of the plurality of first data into the acquired feature quantities ((4) in FIG. 1).

[0020] In this way, the similarity between two feature amounts corresponding to two first data corresponding to two second data belonging to different clusters becomes smaller than the similarity between two feature amounts corresponding to two first data corresponding to two second data belonging to the same cluster when the second data is clustered, and these feature amounts maintain the similarity / dissimilarity relationship of the second data. Therefore, by replacing the first data with feature amounts and anonymizing it, the information processing device 1 can process the first data while taking into account its relationship with the second data.

[0021] [Functional configuration of information processing device 1] Next, a description will be given of the functional configuration of the information processing device 1. FIG.

[0022] As shown in FIG. 2, the information processing device 1 includes a communication unit 11, a storage unit 12, and a control unit 13. The communication unit 11 is a communication interface for transmitting and receiving data to and from the carrier terminal 2 and the like via a communication network.

[0023] The storage unit 12 is a storage medium that stores various types of data, and includes a ROM (Read Only Memory), a RAM (Random Access Memory), a hard disk, an SSD (Solid State Drive), a flash memory, etc. The storage unit 12 stores a program executed by the control unit 13. The storage unit 12 stores a program that causes the control unit 13 to function as an acquisition unit 131, a generation unit 132, a replacement unit 133, and an output unit 134.

[0024] The control unit 13 is, for example, a CPU (Central Processing Unit). The control unit 13 executes a program stored in the storage unit 12, thereby functioning as an acquisition unit 131, a generation unit 132, a replacement unit 133, and an output unit .

[0025] The acquiring unit 131 acquires user data that associates the user IDs of each of a plurality of users collected by a predetermined business operator with first data and second data related to the first data and having a simpler data structure than the first data. The first data is, for example, location history data indicating the user's location history, and the second data is location history-related data created based on the location history data.

[0026] Fig. 3 is a diagram showing an example of user data. As shown in Fig. 3, it can be seen that location history data indicating the user's location history and the number of visits to point A and point B as related data of the location history are associated with the user ID. In the example shown in Fig. 3, for the sake of simplicity, it is assumed that the number of user data items is four.

[0027] The generation unit 132 generates a learning model that outputs, in response to input of first data, feature quantities that indicate the characteristics of the first data. Specifically, the generation unit 132 inputs, into the learning model, two pieces of first data that correspond to two pieces of second data that belong to the same cluster when the second data are clustered, and generates a learning model trained so that the similarity between the two feature quantities output from the learning model is smaller than the similarity indicating the degree of similarity between the two feature quantities output from the learning model. Here, the greater the similarity between the two feature quantities, the more similar the two feature quantities are.

[0028] Specifically, first, the generation unit 132 extracts the user ID and the second data as partial data from the user data acquired by the acquisition unit 131, and performs clustering based on the second data included in the partial data. Here, before extracting the partial data, the generation unit 132 may perform K-anonymization on the second data included in the user data, and perform clustering based on the second data after K-anonymization.

[0029] Fig. 4 shows an example of clustering the second data included in the user data shown in Fig. 3. As shown in Fig. 4, it can be seen that the second data included in the four pieces of user data are classified into two clusters.

[0030] Next, the generation unit 132 generates a learning model that has been trained so that the similarity between two feature quantities output from the learning model by inputting two first data corresponding to two user IDs belonging to different clusters into the learning model is smaller than the similarity between two feature quantities output from the learning model by inputting two first data corresponding to two user IDs belonging to the same cluster in the partial data into the learning model.

[0031] Specifically, P(i) is the index set of users in the same cluster as the i-th user, A(i) is the index set of all users other than the i-th user, and z i , z p , z a are the feature quantities of the i-th, p-th, and a-th users, respectively, and τ is the temperature parameter.

[0032] The generation unit 132 performs learning so that the symmetric loss output by the loss function L shown in the following equation (1) is minimized.

number

[0033] The replacement unit 133 inputs each of the multiple first data included in the user data to the learning model, acquires a feature corresponding to each of the multiple first data, and replaces each of the multiple first data included in the user data with the acquired feature.

[0034] Fig. 5 is a diagram showing an example in which first data included in user data is converted into feature quantities. In the example shown in Fig. 5, feature quantities are expressed as 0 or 1 for ease of explanation, but it can be seen that the location history as the first data is converted into feature quantities. Note that, after converting the first data into feature quantities, the replacement unit 133 may perform K-anonymization on the entire user data to anonymize the user data.

[0035] The output unit 134 outputs the user data obtained by converting the first data. For example, the output unit 134 outputs to the storage unit 12 a file indicating the user data obtained by converting the first data.

[0036] [Operation flow] Next, a description will be given of the flow of processing related to the information processing device 1. Fig. 6 is a flowchart showing the flow of processing in the information processing device 1. First, the acquisition unit 131 acquires user data from the operator terminal 2 (S1).

[0037] Next, the generation unit 132 extracts the user ID and the second data from the user data as partial data (S2), and clusters the partial data based on the second data (S3). Next, the generation unit 132 extracts two user IDs from each of the clustered partial data, and performs training of the learning model using the first data associated with the extracted user IDs (S4). Specifically, the generation unit 132 generates the learning model by training so as to minimize the contrast loss output by the loss function L shown in Equation (1).

[0038] Next, the substitution unit 133 inputs each of the multiple first data included in the user data into the learning model and acquires a feature corresponding to each of the multiple first data (S5).The substitution unit 133 then replaces each of the multiple first data included in the user data with the acquired feature (S6).The output unit 134 outputs the user data after the first data has been converted into the feature (S7).

[0039] [Effects of information processing device 1] As described above, in the information processing system S according to the present embodiment, the information processing device 1 acquires user data that associates user identification information for identifying each of a plurality of users with first data and second data related to the first data, the second data having a simpler data structure than the first data, and generates a learning model that outputs, in response to input of the first data, feature quantities that indicate characteristics of the first data. The learning model inputs two pieces of first data corresponding to two pieces of second data that belong to the same cluster when the second data are clustered, and inputs two pieces of first data corresponding to two pieces of second data that belong to different clusters. The learning model is trained so that the similarity between the two feature quantities output from the learning model is smaller than the similarity indicating the degree of similarity between the two feature quantities. In this way, the information processing device 1 can process the first data while taking into account the relationship between the first data as the original data to be anonymized and the second data as other data.

[0040] Furthermore, this invention can contribute to Goal 9 of the United Nations' Sustainable Development Goals (SDGs), "Build resilient infrastructure, promote inclusive and sustainable industrialization, and promote innovation and infrastructure." All or part of the device can be configured in any unit, functionally or physically, distributed or integrated. New embodiments resulting from any combination of multiple embodiments are also included in the embodiments of the present invention. The effects of the new embodiments resulting from the combination also have the effects of the original embodiments. [Explanation of symbols]

[0041] 1. Information processing equipment 2 1st device 3 Second device 11 Communications Department 12 Storage section 13 Control Unit 131 Acquisition Department 132 Generation part 133 Substitution part 134 Output section

Claims

1. an acquisition unit that acquires user data that associates user identification information for identifying each of a plurality of users collected by a predetermined business operator, first data, and second data related to the first data, the second data having a simpler data structure than the first data; a generation unit that generates a learning model that outputs, in response to input of the first data, feature quantities that indicate features of the first data, the learning model inputting two pieces of first data that correspond to two pieces of second data that belong to the same cluster when the second data are clustered, into the learning model, inputting two pieces of first data that correspond to two pieces of second data that belong to different clusters into the learning model, and generating a learning model that has been trained so that the similarity between the two feature quantities output from the learning model is smaller than the similarity indicating the degree of similarity between the two feature quantities output from the learning model; a replacement unit that inputs each of a plurality of first data included in the user data into the generated learning model, acquires feature amounts corresponding to each of the plurality of first data, and replaces each of the plurality of first data included in the user data with the acquired feature amounts; An information processing device having the above.

2. the generation unit performs K-anonymization on the second data, inputs into the learning model two pieces of first data corresponding to two pieces of second data belonging to the same cluster when the second data after K-anonymization is clustered, inputs into the learning model first pieces of first data corresponding to two pieces of second data belonging to different clusters, and generates the learning model trained so that the similarity between the two feature amounts output from the learning model is smaller than the similarity between the two feature amounts output from the learning model; The information processing device according to claim 1 .

3. the generation unit extracts the user identification information and the second data as partial data from the user data acquired by the acquisition unit, clusters the partial data, and generates the learning model trained so that the similarity between two feature amounts output from the learning model by inputting two pieces of first data corresponding to two pieces of user identification information belonging to different clusters into the learning model is smaller than the similarity between two feature amounts output from the learning model by inputting two pieces of first data corresponding to two pieces of user identification information belonging to the same cluster into the learning model; The information processing device according to claim 1 .

4. The first data is location history data indicating a location history of the user, and the second data is data related to the location history created based on the location history data. The information processing device according to claim 1 .

5. The computer executes acquiring user data that associates user identification information for identifying each of a plurality of users collected by a predetermined business operator, first data, and second data related to the first data, the second data having a simpler data structure than the first data; a learning model that outputs a feature quantity indicating a feature of the first data in response to input of the first data, the learning model inputting two pieces of first data corresponding to two pieces of second data that belong to the same cluster when the second data is clustered, and inputting two pieces of first data corresponding to two pieces of second data that belong to different clusters into the learning model, and generating a learning model trained so that the similarity between the two feature quantities output from the learning model is smaller than the similarity indicating the degree of similarity between the two feature quantities output from the learning model; inputting each of a plurality of first data included in the user data into the generated learning model, acquiring feature amounts corresponding to each of the plurality of first data, and replacing each of the plurality of first data included in the user data with the acquired feature amounts; An information processing method comprising:

Citation Information

Patent Citations

  • Coordination server program, business operator server program, and data coordinated system

    JP2021117679A

  • Training data generation apparatus, learning model generation apparatus, and method of generating training data

    JP2023013293A

  • Information processing device, information processing method, program, and information processing system

    JP2023117614A

  • Search device, search method, and recording medium

    WO2022091299A1