Information processing device and information processing method
The information processing device clusters and combines anonymized user data using machine learning, addressing the challenge of combining anonymized data for analysis and enabling personalized information delivery.
Patent Information
- Application Number
- JP2024033800
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-06
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-03-06
AI Technical Summary
Anonymized user information from multiple businesses cannot be combined, preventing effective data analysis.
An information processing device that clusters anonymized user information from different businesses using machine learning models to generate clusters, acquires parameters for each cluster, and combines them based on these parameters to enable data analysis while maintaining user anonymity.
Enables the combination of anonymized user information across businesses, facilitating effective data analysis and personalized information delivery to users.
Smart Images

Figure 2025135815000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device and an information processing method. [Background technology]
[0002] Conventionally, user information, which is information about users, is collected from multiple businesses, and the collected information is combined to perform data analysis. For example, Patent Document 1 discloses a system that combines data about personal information of users corresponding to multiple businesses based on a combination key for combining multiple pieces of user information. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2021-117679 Summary of the Invention [Problem to be solved by the invention]
[0004] In order to protect users' personal information, user information is anonymized. However, if multiple businesses each anonymize their user information, this user information cannot be combined, which creates a problem in that data analysis cannot be performed.
[0005] The present invention has been made in consideration of these points, and has as its object to make it possible to combine information about a plurality of anonymized users. [Means for solving the problem]
[0006] An information processing device according to a first aspect of the present invention includes a first acquisition unit that acquires a first information group including first attribute information that is attribute information indicating a user's attribute and collected by a first business operator, and a second information group that is attribute information that is collected by a second business operator different from the first business operator and includes second attribute information that is attribute information indicating a user's attribute; a generation unit that performs clustering of the attribute information based on the first attribute information included in the first information group to generate a plurality of first clusters that are clusters corresponding to the first information group, and that performs clustering of the attribute information based on the second attribute information included in the second information group to generate a plurality of second clusters that are clusters corresponding to the second information group; and a parameter acquisition unit that acquires a plurality of clusters in response to input of attribute information. a second acquisition unit that inputs the attribute information corresponding to each of the plurality of first clusters into a machine learning model that outputs a parameter set, and acquires first parameters that are the parameters of each of the plurality of first clusters from the machine learning model, and inputs the attribute information corresponding to each of the plurality of second clusters into the machine learning model and acquires second parameters that are the parameters corresponding to each of the plurality of second clusters from the machine learning model; and a combining unit that combines one or more second clusters with at least one of the plurality of first clusters based on the first parameters of each of the plurality of first clusters and the second parameters of each of the plurality of second clusters.
[0007] The second acquisition unit may acquire one or more first parameters corresponding to each of the plurality of first clusters, and may acquire one or more second parameters corresponding to each of the plurality of second clusters, using the machine learning model that outputs one or more parameters indicating characteristics of a person having the attribute information in response to input of the attribute information.
[0008] The generation unit may generate the first cluster and the second cluster associated with range information indicating a range of attribute values summarized by clustering, and the second acquisition unit may acquire the first parameters from the machine learning model by inputting the attributes corresponding to each of the plurality of first clusters and the range information corresponding to the attributes to the machine learning model, which outputs parameters for combining a plurality of clusters in response to input of the attributes and the range information corresponding to the attributes, and may acquire the second parameters from the machine learning model by inputting the attributes corresponding to each of the plurality of second clusters and the range information corresponding to the attributes to the machine learning model.
[0009] The information processing device may further include a distribution unit that distributes information corresponding to the attribute information indicated by the second cluster combined with the first cluster to a terminal of a user having the attribute information indicated by the first cluster.
[0010] The generation unit may calculate a first cluster membership degree indicating the degree to which a user corresponding to each of a plurality of records belongs to the first cluster based on the degree of match between attribute information indicated by the generated first cluster and attribute information of each of a plurality of records included in the first information group, and the distribution unit may preferentially distribute information corresponding to the attribute information indicated by the second cluster combined with the first cluster to terminals of users whose first cluster membership degree to the first cluster is relatively high.
[0011] The first acquisition unit may acquire the first information group from a first device used by the first operator and acquire the second information group from a second device used by the second operator, and the information processing device may further have a transmission unit that transmits information indicating the combination of the first cluster and the second cluster combined by the combination unit to at least one of the first device and the second device.
[0012] An information processing method according to a second aspect of the present invention is executed by a computer, and includes the steps of acquiring a first information group including first attribute information that is attribute information indicating user attributes and collected by a first business operator, and a second information group including second attribute information that is attribute information indicating user attributes and collected by a second business operator different from the first business operator, clustering the attribute information based on the first attribute information included in the first information group to generate a plurality of first clusters that are clusters corresponding to the first information group, and clustering the attribute information based on the second attribute information included in the second information group to generate a plurality of second clusters that are clusters corresponding to the second information group, and combining the plurality of clusters in response to input attribute information. inputting the attribute information corresponding to each of the plurality of first clusters into a machine learning model that outputs parameters of the plurality of first clusters, and acquiring first parameters that are the parameters of each of the plurality of first clusters from the machine learning model, and inputting the attribute information corresponding to each of the plurality of second clusters into the machine learning model, and acquiring second parameters that are the parameters corresponding to each of the plurality of second clusters from the machine learning model; and combining one or more of the second clusters with at least one of the plurality of first clusters based on the first parameters of each of the plurality of first clusters and the second parameters of each of the plurality of second clusters. [Effects of the Invention]
[0013] According to the present invention, it is possible to combine a plurality of pieces of anonymous user information. [Brief explanation of the drawings]
[0014] [Figure 1] FIG. 1 is a diagram illustrating an overview of an information processing system. [Figure 2] FIG. 2 is a diagram illustrating a functional configuration of an information processing device. [Figure 3] 4A and 4B are diagrams showing examples of a first information group and a second information group. [Figure 4]4 is a diagram showing an example of a first cluster generated from the first information group shown in FIG. 3. FIG. [Figure 5] 4 is a diagram showing an example of a second cluster generated from the second information group shown in FIG. 3. FIG. [Figure 6] FIG. 10 is a diagram showing an example of a template of instruction information. [Figure 7] FIG. 10 is a diagram showing an example of instruction information including attribute information corresponding to a first cluster. [Figure 8] 8 is a diagram showing an example of text information acquired from a large-scale language model in response to instruction information corresponding to FIG. 7. FIG. [Figure 9] FIG. 4 is a diagram illustrating an example of a first parameter and a second parameter. [Figure 10] FIG. 10 is a diagram illustrating an example of coupled cluster information. [Figure 11] 10 is a flowchart showing a processing flow in the information processing device. [Figure 12] FIG. 1 is a diagram illustrating an overview of an information processing system in which a first device and a second device form a cluster. DETAILED DESCRIPTION OF THE INVENTION
[0015] [Outline of Information Processing System S] 1 is a diagram illustrating an overview of an information processing system S. The information processing system S includes an information processing device 1, a first device 2, and a second device 3, and is a system that combines user information, which is information about users of each of a plurality of businesses.
[0016] The information processing device 1 is, for example, a computer used by an aggregator that aggregates data and provides a service of providing the aggregated data. The information processing device 1 is communicably connected to a first device 2 used by a first operator and a second device 3 used by a second operator different from the first operator via a communication network (not shown) such as the Internet or a mobile phone line.
[0017] The information processing device 1 acquires from the first device 2 a first information group including first attribute information, which is attribute information indicating user attributes collected by a first business operator, and acquires from the second device 3 a second information group including second attribute information, which is attribute information indicating user attributes collected by a second business operator ((1) and (2) in Figure 1).
[0018] The information processing device 1 clusters the attribute information based on the first attribute information included in the first information group, and generates a plurality of first clusters that correspond to the first information group. Similarly, the information processing device 1 clusters the attribute information based on the second attribute information included in the second information group, and generates a plurality of second clusters that correspond to the second information group ((3) in FIG. 1).
[0019] The information processing device 1 inputs attribute information corresponding to each cluster to a machine learning model that outputs parameters for combining multiple clusters in response to input attribute information ((4) in FIG. 1), and acquires parameters corresponding to each cluster from the machine learning model ((5) in FIG. 1). Specifically, the information processing device 1 inputs attribute information corresponding to each of multiple first clusters to the machine learning model, and acquires first parameters that are parameters for each of the multiple first clusters from the machine learning model. Similarly, the information processing device 1 inputs attribute information corresponding to each of multiple second clusters to the machine learning model, and acquires second parameters that are parameters corresponding to each of the multiple second clusters from the machine learning model.
[0020] The information processing device 1 combines one or more second clusters with at least one of the plurality of first clusters based on the first parameter of each of the plurality of first clusters and the second parameter of each of the plurality of second clusters ((6) in FIG. 1). In this way, the information processing device 1 can combine a plurality of pieces of anonymized user information.
[0021] [Functional configuration of information processing device 1] Next, a description will be given of the functional configuration of the information processing device 1. FIG.
[0022] As shown in FIG. 2, the information processing device 1 includes a communication unit 11, a storage unit 12, and a control unit 13. The communication unit 11 is a communication interface for transmitting and receiving data to and from the first device 2, the second device 3, and the like via a communication network.
[0023] The storage unit 12 is a storage medium that stores various types of data, and includes a read-only memory (ROM), a random access memory (RAM), a hard disk, a solid-state drive (SSD), a flash memory, etc. The storage unit 12 stores programs executed by the control unit 13. The storage unit 12 stores programs that cause the control unit 13 to function as a first acquisition unit 131, a generation unit 132, a second acquisition unit 133, a combination unit 134, and a transmission unit 135.
[0024] The control unit 13 is, for example, a CPU (Central Processing Unit). The control unit 13 executes a program stored in the storage unit 12, thereby functioning as a first acquisition unit 131, a generation unit 132, a second acquisition unit 133, a combination unit 134, and a transmission unit 135.
[0025] The first acquisition unit 131 acquires a first information group including first attribute information, which is attribute information indicating user attributes collected by a first business operator, and a second information group including second attribute information, which is attribute information indicating user attributes collected by a second business operator different from the first business operator. For example, the first acquisition unit 131 acquires the first information group from a first device 2 used by the first business operator, and acquires the second information group from a second device 3 used by the second business operator.
[0026] FIG. 3 is a diagram illustrating an example of a first information group and a second information group. FIG. 3(a) illustrates the first information group, and FIG. 3(b) illustrates the second information group. As illustrated in FIG. 3(a), the first information group includes an anonymized user ID as user identification information for identifying a user. The first attribute information associates the user ID with two attributes: a static attribute "user generation" and a dynamic attribute that may change over time, indicating the model of the smartphone used by the user. As illustrated in FIG. 3(b), the second information group includes an anonymized user ID for identifying a user. The second information group also associates the user ID with two dynamic attributes: a dynamic attribute "purchased product category" indicating the category of products purchased by the user, and a dynamic attribute "point usage" indicating the amount of points used by the user for a purchase.
[0027] Although the first information group and the second information group each include an anonymized user ID, this is not limiting. For example, the first information group and the second information group may not include a user ID, or may include a user ID in a different format that cannot be associated with the user ID.
[0028] The generation unit 132 clusters the attribute information based on the first attribute information included in the first information group to generate one or more first clusters that are clusters corresponding to the first information group, and clusters the attribute information based on the second attribute information included in the second information group to generate one or more second clusters that are clusters corresponding to the second information group. The generation unit 132 generates a plurality of first clusters and second clusters associated with range information that indicates ranges of values indicated by the attributes grouped together by clustering.
[0029] For example, the generation unit 132 uses a clustering method such as the k-means method or LDA (Latent Dirichlet Allocation) to generate multiple first clusters that correspond to the first information group, and generate multiple second clusters that correspond to the second information group.
[0030] FIG. 4 is a diagram showing an example of a first cluster generated from the first information group shown in FIG. 3. FIG. 4(a) shows an example in which a first cluster ID for identifying a first cluster is assigned to each record included in the first information group. FIG. 4(b) shows first cluster information showing the contents of the first cluster. The first cluster information is, for example, information in which the first cluster IDs of multiple first clusters are associated with range information showing the range of attribute values corresponding to each of multiple attributes corresponding to the first cluster. For example, as a result of clustering a record with a first cluster ID of "1," it can be confirmed that "young" is identified as the range information corresponding to the attribute "generation" included in the first attribute information, and "latest model" is identified as the range information corresponding to the attribute "model used."
[0031] After generating the first cluster, the generation unit 132 may calculate a degree of match between attribute information indicated by the generated first cluster and attribute information of each of the multiple records included in the first information group. Then, based on the calculated degree of match, the generation unit 132 may calculate a first cluster membership degree indicating the degree to which a user corresponding to each of the multiple records belongs to the first cluster. This calculation can be performed using a clustering method such as the well-known LDA (Latent Dirichlet Allocation).
[0032] FIG. 5 is a diagram showing an example of a second cluster generated from the second information group shown in FIG. 3. FIG. 5(a) shows an example in which a second cluster ID for identifying a second cluster is assigned to each record included in the second information group. FIG. 4(b) shows second cluster information showing the contents of a second cluster. The second cluster information is, for example, information in which the second cluster IDs of multiple second clusters are associated with range information indicating the range of attribute values corresponding to each of multiple attributes. For example, as a result of clustering a record with a second cluster ID of "1," it can be seen that "health / outdoors" is identified as the range information corresponding to the attribute "purchase product category" included in the second attribute information, and "few" is identified as the range information for the attribute "points usage."
[0033] Note that the generation unit 132 may generate clusters so that the number of data constituting the cluster, that is, the number of users constituting the cluster, satisfies k-anonymity. In this way, the information processing device 1 can perform clustering while ensuring the anonymity of users.
[0034] The second acquisition unit 133 inputs attribute information corresponding to each of a plurality of first clusters to a machine learning model that outputs combining parameters, which are parameters for combining a plurality of clusters in response to input of attribute information, and acquires the first parameters, which are combining parameters for each of the plurality of first clusters, from the machine learning model. Also, the second acquisition unit 133 inputs attribute information corresponding to each of a plurality of second clusters to the machine learning model and acquires the second parameters, which are combining parameters, corresponding to each of the plurality of second clusters from the machine learning model.
[0035] Here, the machine learning model is, for example, generative AI (Artificial Intelligence) such as large-scale language models (LLMs) that output text information containing parameters for combining multiple clusters in response to input attribute information. The large-scale language model can output one or more combining parameters that indicate the characteristics of a person who has the attribute information in response to input of attributes as attribute information and range information corresponding to the attributes.
[0036] First, the second acquisition unit 133 generates instruction information including attribute information. For example, a template of the instruction information is stored in the storage unit 12. FIG. 6 is a diagram showing an example of the template of the instruction information. In the example shown in FIG. 6, as user information, "@propaerty1" and "@propaerty2" are shown as parameters for specifying attributes, and "@values1" and "@values2" are shown as parameters for specifying range information corresponding to the attributes, of which "@value1" and "@value2" are shown to be taken.
[0037] The second acquisition unit 133 generates instruction information to be input to the large-scale language model by reflecting attributes corresponding to the clusters and range information corresponding to the attributes in a template of instruction information stored in the storage unit 12. FIG. 7 is a diagram showing an example of instruction information including attribute information corresponding to the first cluster. From the example shown in FIG. 7, it can be seen that the parameters shown in the template of instruction information shown in FIG. 6 reflect the attributes of the first cluster having a first cluster ID of "1," such as "age" and "model used," as well as the range information corresponding to the attributes, such as "young" out of "young" and "middle-aged and elderly," and "latest model" out of "latest model" and "older model."
[0038] Next, the second acquisition unit 133 inputs, to the large-scale language model, instruction information including attributes corresponding to each of the multiple first clusters and range information corresponding to the attributes, and then acquires, from the large-scale language model, text information including one or more first parameters corresponding to each of the multiple first clusters.
[0039] Similarly, the second acquisition unit 133 inputs instruction information including attributes corresponding to each of the multiple second clusters and range information corresponding to the attributes to the large-scale language model, and acquires text information including one or more second parameters corresponding to each of the multiple second clusters from the large-scale language model.
[0040] Fig. 8 is a diagram showing an example of text information acquired from a large-scale language model for the instruction information corresponding to Fig. 7. Fig. 8 shows that the first cluster contains values "3," "2," and "4" corresponding to the items "attitude," "subjective norm," and "possibility of behavioral control," which indicate the characteristics of the person. These values are used as the first parameters.
[0041] It is desirable to specify the range of values that the parameter can take for each item of the first parameter, such as "on a 7-point scale" as shown in Figure 7. It is also desirable to specify more specifically, "7-point scale (1: not applicable at all, 2: not applicable somewhat, 3: not applicable to some extent, 4: neither applicable nor unapplicable, 5: applicable to some extent, 6: applicable somewhat, 7: applicable to all extent)."
[0042] Next, the second acquisition unit 133 analyzes text information acquired from the large-scale language model in response to the instruction information corresponding to the first cluster, and extracts, as first parameters, numerical values corresponding to the items "attitude," "subjective norm," and "possibility of behavioral control" contained in the text information, which indicate the characteristics of the person corresponding to the first cluster. Similarly, the second acquisition unit 133 analyzes text information acquired from the large-scale language model in response to the instruction information corresponding to the second cluster, and extracts, as second parameters, numerical values corresponding to the items "attitude," "subjective norm," and "possibility of behavioral control" contained in the text information, which are the characteristics of the person corresponding to the second cluster and have the same indicators as the first parameters, thereby acquiring second parameters.
[0043] FIG. 9 is a diagram showing an example of a first parameter and a second parameter. FIG. 9(a) shows the first parameter acquired for the first cluster shown in FIG. 4, and FIG. 9(b) shows the second parameter acquired for the second cluster shown in FIG. 5. As shown in FIGS. 9(a) and 9(b), it can be seen that values corresponding to the items "attitude," "subjective norm," and "possibility of behavioral control," which indicate a person's characteristics, are acquired as the first parameter and the second parameter. Note that, although the items indicating a person's characteristics are described as "attitude," "subjective norm," and "possibility of behavioral control," they are not limited to these, and other items indicating a person's characteristics may also be used.
[0044] The second acquisition unit 133 may generate multiple first parameters for each of the multiple first clusters. The second acquisition unit 133 may generate multiple first parameters by changing a temperature parameter of the large-scale language model to output multiple parameters for the same input. Furthermore, if multiple combinations of attribute information correspond to the first cluster, the second acquisition unit 133 may generate multiple first parameters by causing the large-scale language model to generate multiple first parameters based on each combination. For example, if the first cluster is a young age cluster and includes both the latest model and an older model, the second acquisition unit 133 inputs both "young age" and "latest model" and "young age" and "old model" into the large-scale language model and generates first parameters corresponding to each combination. Such a parameter generation method can also be applied to the second parameters of each second cluster.
[0045] The combining unit 134 combines one or more second clusters with at least one of the multiple first clusters based on the first parameters of each of the multiple first clusters and the second parameters of each of the multiple second clusters acquired by the second acquiring unit 133. For example, the combining unit 134 calculates parameter similarities, such as cosine similarities, between the multiple first parameters of each of the multiple first clusters and the multiple second parameters of each of the multiple second clusters.
[0046] Then, the combining unit 134 combines the first cluster whose calculated similarity exceeds a predetermined threshold with the second cluster. For example, the combining unit 134 generates combined cluster information by combining the first cluster whose calculated similarity exceeds a predetermined threshold with the second cluster. Here, the combining unit 134 may combine, with one first cluster, multiple second clusters having second parameters whose similarity to the first parameter of the first cluster exceeds a predetermined threshold.
[0047] Note that combining unit 134 combines the first cluster and the second cluster based on the similarity of a parameter such as cosine similarity, but this is not limiting. For example, the distance from the first cluster to the second cluster may be regarded as the transportation cost, and the combination of the first cluster and the second cluster that minimizes the transportation cost from one or more first clusters to one or more second clusters using the Synchorn algorithm may be determined as the first cluster and the second cluster to be combined.
[0048] Fig. 10 is a diagram showing an example of joined cluster information. From Fig. 10, it can be seen that a first cluster with a first cluster ID of "1" is joined to a second cluster with a second cluster ID of "2", and that a first cluster with a first cluster ID of "2" is joined to a second cluster with a second cluster ID of "1".
[0049] The transmitting unit 135 transmits information indicating the combination of the first cluster and the second cluster combined by the combining unit 134 to at least one of the first device 2 and the second device 3. For example, the transmitting unit 135 distributes the combined cluster information shown in FIG. 10 as information indicating the combination of the first cluster and the second cluster combined by the combining unit 134 to at least one of the first device 2 and the second device 3.
[0050] The transmitting unit 135 may function as a distribution unit and distribute information corresponding to the attribute information indicated by the second cluster combined with the first cluster to a terminal of a user having the attribute information indicated by the first cluster. In this case, for example, the transmitting unit 135 selects information corresponding to the value of the second attribute information associated with the first cluster ID for a user with the first cluster ID. For example, in the example shown in FIG. 10, the transmitting unit 135 selects, as advertising information to be distributed to a user with a first cluster ID of "2," advertising information corresponding to "health and outdoor activities," which is attribute information corresponding to the second cluster ID "1" associated with the first cluster ID.
[0051] Then, the transmitting unit 135 associates the user ID associated with the first cluster ID with the selected advertising information and transmits them to the first device 2, thereby causing the first device 2 to deliver the advertising information to the terminal used by the user of the user ID. For example, when the transmitting unit 135 selects advertising information corresponding to "health and outdoor" as advertising information to be delivered to the user whose first cluster ID is "2", the transmitting unit 135 transmits instruction information to the first device 2 instructing the first device 2 to deliver the advertising information to the terminal of the user whose user ID is associated with the first cluster ID "2" shown in Fig. 4(a). In this way, the information processing device 1 can deliver effective information to the user.
[0052] Furthermore, although transmitting unit 135 distributes information corresponding to the attribute information indicated by the second cluster combined with the first cluster to a terminal of a user having the attribute information indicated by the first cluster, this is not limiting. For example, transmitting unit 135 may distribute information corresponding to the attribute information indicated by the second cluster combined with the first cluster to a terminal of a user not having the attribute information indicated by the first cluster.
[0053] For example, when the generation unit 132 calculates the first cluster belonging degree for each user corresponding to each first information group, the transmission unit 135 identifies users who have a relatively high first cluster belonging degree for the first cluster from among users who do not have attribute information indicated by the first cluster. Then, the transmission unit 135 may preferentially deliver information corresponding to the attribute information indicated by the second cluster combined with the first cluster to terminals of the identified users who have a relatively high first cluster belonging degree.
[0054] For example, the transmission unit 135 may associate the user ID of the identified user having a relatively high degree of first cluster belongingness, which is included in the first information group, with the selected advertising information and transmit the associated information to the first device 2, thereby causing the first device 2 to deliver the advertising information to a terminal used by the user of the user ID. If a budgetary limit for information delivery is set, the transmission unit 135 may preferentially identify users having a high degree of first cluster belongingness within the budgetary limit.
[0055] [Operation flow] Next, a description will be given of the flow of processing related to the information processing device 1. Fig. 11 is a flowchart showing the flow of processing in the information processing device 1. First, the first acquisition unit 131 acquires a first information group from the first device 2 and acquires a second information group from the second device 3 (S1).
[0056] Next, the generation unit 132 clusters the attribute information based on the first attribute information included in the acquired first information group to generate multiple first clusters, and also clusters the attribute information based on the second attribute information included in the second information group to generate multiple second clusters (S2).
[0057] Next, the second acquiring unit 133 inputs attribute information corresponding to each of the plurality of first clusters into the machine learning model, and acquires text information including first parameters that are parameters for combining each of the plurality of first clusters. Then, the second acquiring unit 133 acquires the first parameters included in the acquired text information (S3).
[0058] Next, the second acquisition unit 133 inputs attribute information corresponding to each of the second clusters to the machine learning model, and acquires text information including second parameters that are combining parameters for each of the second clusters. Then, the second acquisition unit 133 acquires the second parameters included in the acquired text information (S4).
[0059] Next, the combining unit 134 generates combined cluster information by combining one or more second clusters with at least one of the multiple first clusters based on the first parameter of each of the multiple first clusters and the second parameter of each of the multiple second clusters acquired by the second acquiring unit 133 (S5). Next, the transmitting unit 135 transmits the combined cluster information generated in S7 to at least one of the first device 2 and the second device 3 (S6).
[0060] [Variation 1] In the above-described embodiment, the information processing device 1 generates the first cluster and the second cluster, but this is not limiting. The devices that generate the first cluster and the second cluster may be the first device 2 and the second device 3, respectively. Fig. 12 is a diagram illustrating an overview of the information processing system S in the case where the first device 2 and the second device 3 generate clusters.
[0061] 12, the first device 2 generates a first cluster based on the first information group ((1) in FIG. 12) and transmits attribute information of the first cluster to the information processing device 1 ((2) in FIG. 12). The second device 3 generates a second cluster based on the second information group ((3) in FIG. 12) and transmits attribute information of the second cluster to the information processing device 1 ((4) in FIG. 12).
[0062] The information processing device 1 inputs attribute information corresponding to each cluster obtained from the first device 2 and the second device 3 into the machine learning model ((5) in Figure 12), and obtains parameters corresponding to each cluster from the machine learning model ((6) in Figure 12).
[0063] The information processing device 1 combines one or more second clusters with at least one of the multiple first clusters based on the first parameter of each of the multiple first clusters and the second parameter of each of the multiple second clusters ((7) in FIG. 12).
[0064] In this case, the first device 2 may hold association information between a first cluster ID and a user ID, and the second device 3 may hold association information between a second cluster ID and a user ID. Then, the transmission unit 135 of the information processing device 1 may transmit advertising information selected corresponding to the first cluster combined with the second cluster to the first device 2 in association with the first cluster ID of the first cluster, and the first device 2 may deliver the advertising information to a terminal used by a user whose user ID is associated with the first cluster ID. Similarly, the transmission unit 135 of the information processing device 1 may transmit advertising information selected corresponding to the second cluster combined with the first cluster to the second device 3 in association with the second cluster ID of the second cluster, and the second device 3 may deliver the advertising information to a terminal used by a user whose user ID is associated with the second cluster ID.
[0065] Furthermore, when the information processing device 1 generates the first cluster and the second cluster, it may provide dedicated sections for generating the first cluster and the second cluster so that the input data and processing process cannot be confirmed by each other. In this case, the transmission unit 135 may associate the cluster ID of the first cluster with a user ID in the section corresponding to the first cluster, and distribute the advertising information selected corresponding to the first cluster to the terminal used by the user of the user ID. In this way, security and anonymity can be ensured when handling the first information group and the second information group.
[0066] [Variation 2] Furthermore, in the above-described embodiment, the information processing device 1 acquires each parameter by inputting text information into a large-scale language model. However, the information input to the machine learning model is not limited to text information, and information other than large-scale language models that perform image and video processing, text analysis, and speech recognition can also be used. For example, a user's hobbies and preferences may be read from the wallpaper, etc., set by the user and reflected in the parameters. Furthermore, if media image analysis is performed using the user's voice information, which is now widely common, such as predicting behavior, various data can be used to acquire parameters, although user consent is required.
[0067] [Effects of information processing device 1] As described above, the information processing device 1 according to the present embodiment acquires a first information group including first attribute information indicating user attributes collected by a first service provider and a second information group including second attribute information indicating user attributes collected by a second service provider, generates multiple first clusters corresponding to the first information group, and generates multiple second clusters corresponding to the second information group. The information processing device 1 then inputs attribute information corresponding to each of the multiple first clusters into a machine learning model, acquires first parameters for each of the multiple first clusters from the machine learning model, inputs attribute information corresponding to each of the multiple second clusters into the machine learning model, acquires second parameters for each of the multiple second clusters from the machine learning model, and combines one or more second clusters with at least one of the multiple first clusters based on the first parameters for each of the multiple first clusters and the second parameters for each of the multiple second clusters. In this way, the information processing device 1 can combine multiple pieces of anonymized user information.
[0068] Furthermore, this invention will make it possible to contribute to Goal 9 of the United Nations' Sustainable Development Goals (SDGs), which is "Build resilient infrastructure, promote inclusive and sustainable industrialization, and promote innovation and resilience."
[0069] The present invention has been described above using embodiments, but the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. For example, all or part of the device can be configured by functionally or physically distributing or integrating any unit. Furthermore, new embodiments resulting from any combination of multiple embodiments are also included in the embodiments of the present invention. The effects of the new embodiments resulting from the combination also have the effects of the original embodiments. [Explanation of symbols]
[0070] 1. Information processing equipment 2 1st device 3 Second device 11 Communications Department 12 Storage section 13 Control Unit 131 First acquisition part 132 Generation part 133 Second Acquisition Department 134 Joint 135 Transmitter
Claims
1. a first acquisition unit that acquires a first information group including first attribute information that is attribute information indicating an attribute of a user and that is collected by a first business operator, and a second information group that includes second attribute information that is attribute information indicating an attribute of a user and that is collected by a second business operator different from the first business operator; a generation unit that performs clustering of attribute information based on first attribute information included in the first information group to generate a plurality of first clusters that are clusters corresponding to the first information group, and that performs clustering of attribute information based on the second attribute information included in the second information group to generate a plurality of second clusters that are clusters corresponding to the second information group; a second acquisition unit that inputs attribute information corresponding to each of the plurality of first clusters into a machine learning model that outputs parameters for combining a plurality of clusters in response to input of attribute information, and acquires first parameters that are the parameters for each of the plurality of first clusters from the machine learning model, and inputs attribute information corresponding to each of the plurality of second clusters into the machine learning model, and acquires second parameters that are the parameters corresponding to each of the plurality of second clusters from the machine learning model; a combining unit that combines one or more of the second clusters with at least one of the plurality of first clusters based on the first parameter of each of the plurality of first clusters and the second parameter of each of the plurality of second clusters; An information processing device having the above.
2. the second acquisition unit acquires one or more first parameters corresponding to each of the plurality of first clusters and one or more second parameters corresponding to each of the plurality of second clusters, using the machine learning model that outputs one or more parameters indicating characteristics of a person having the attribute information in response to input of the attribute information; The information processing device according to claim 1 .
3. the generation unit generates the first cluster and the second cluster associated with range information indicating a range of values of attributes grouped by clustering; the second acquisition unit acquires the first parameters from the machine learning model by inputting the attributes corresponding to each of the plurality of first clusters and the range information corresponding to the attributes to the machine learning model, which outputs parameters for combining a plurality of clusters in response to input of the attributes and the range information corresponding to the attributes, and acquires the second parameters from the machine learning model by inputting the attributes corresponding to each of the plurality of second clusters and the range information corresponding to the attributes to the machine learning model; The information processing device according to claim 1 .
4. a distribution unit configured to distribute information corresponding to the attribute information indicated by the second cluster combined with the first cluster to a terminal of a user having the attribute information indicated by the first cluster; The information processing device according to claim 1 .
5. the generation unit calculates a first cluster belonging degree indicating a degree to which a user corresponding to each of a plurality of records belongs to the first cluster based on a degree of match between attribute information indicated by the generated first cluster and attribute information of each of a plurality of records included in the first information group; the distribution unit preferentially distributes information corresponding to attribute information indicated by the second cluster combined with the first cluster to a terminal of a user having a relatively high degree of first cluster belongingness to the first cluster. The information processing device according to claim 4 .
6. the first acquisition unit acquires the first information group from a first device used by the first business operator and acquires a second information group from a second device used by the second business operator; a transmitting unit configured to transmit information indicating the combination of the first cluster and the second cluster combined by the combining unit to at least one of the first device and the second device; The information processing device according to claim 1 .
7. The computer executes acquiring a first information group including first attribute information, which is attribute information indicating an attribute of a user and collected by a first business operator, and a second information group including second attribute information, which is attribute information indicating an attribute of a user and collected by a second business operator different from the first business operator; clustering attribute information based on first attribute information included in the first information group to generate a plurality of first clusters that are clusters corresponding to the first information group, and clustering attribute information based on the second attribute information included in the second information group to generate a plurality of second clusters that are clusters corresponding to the second information group; inputting attribute information corresponding to each of the plurality of first clusters into a machine learning model that outputs parameters for combining a plurality of clusters in response to input of attribute information, and acquiring first parameters that are the parameters for each of the plurality of first clusters from the machine learning model; inputting attribute information corresponding to each of the plurality of second clusters into the machine learning model, and acquiring second parameters that are the parameters corresponding to each of the plurality of second clusters from the machine learning model; combining one or more second clusters with at least one of the first clusters based on the first parameter of each of the first clusters and the second parameter of each of the second clusters; An information processing method comprising:
Citation Information
Patent Citations
Coordination server program, business operator server program, and data coordinated system
JP2021117679A