Analysis system, computer implementing method, and computer program
The analysis system addresses the limitation of existing methods by generating networks with vertices corresponding to elements, enabling the analysis of individual characteristics and network features, including the impact of personal attributes on interactions.
Patent Information
- Application Number
- JP2024027943
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-27
- Publication Date
- 2025-09-08
AI Technical Summary
Existing methods for analyzing network features, such as those described in Non-Patent Document 1, cannot distinguish between vertices with the same degree, limiting the ability to analyze individuals within a group represented by the network.
An analysis system that includes a processor for acquiring data on the number of interactions among elements, randomly generating networks with vertices corresponding to these elements, and performing analysis based on these networks, allowing for the distinction between vertices and the analysis of individual characteristics.
Enables the analysis of individual characteristics within a network, including the probability of belonging to a cluster, importance of individuals, and the impact of features like gender and age on interpersonal interactions, while accurately estimating network characteristics.
Smart Images

Figure 2025130633000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an analysis system, a computer-implemented method, and a computer program. [Background technology]
[0002] Non-Patent Document 1 discloses that the results of a survey on the number of interactions of each member of a community are treated as a degree distribution of the vertices of a network, a network that conforms to such a degree distribution is randomly generated, and network features are estimated. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2008-210363 [Non-patent literature]
[0004] [Non-Patent Document 1] Xu Baige et al., "On Estimating Network Features from Interaction Survey Data," Proceedings of the 2022 Annual Meeting of the Japan Society for Industrial and Applied Mathematics, Japan Society for Industrial and Applied Mathematics, 2022 Summary of the Invention
[0005] The technology described in Non-Patent Document 1 can analyze the global characteristics of the network as a whole (the characteristics of the entire group represented by the network), but has the problem that it cannot analyze the individuals belonging to the group represented by the network. This is because multiple networks randomly generated to conform to the degree distribution are simply generated so as not to contradict the degree distribution, and it is not possible to distinguish between vertices with the same degree. Note that this problem can also occur when the elements belonging to the group represented by the network are other than individuals.
[0006] Therefore, it is desirable to solve this problem.
[0007] One aspect of the present disclosure is an analysis system. The disclosed analysis system includes a processor that executes operations including acquiring analysis target data, which is data related to a plurality of elements, randomly generating a plurality of networks, and performing analysis based on the plurality of networks. The analysis target data includes the number of other elements associated with each of the plurality of elements. The number of other elements indicates the number of other elements associated with each of the plurality of elements. Each of the plurality of randomly generated networks has vertices in one-to-one correspondence with the plurality of elements. Each vertex has a number of branches corresponding to the number of other elements associated with the element corresponding to the vertex.
[0008] Other aspects of the present disclosure are computer-implemented methods and computer programs.
[0009] Further details will be described in the following embodiments. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a diagram showing the configuration of the analysis system. [Figure 2] FIG. 2 is a flowchart showing the analysis processing procedure. [Figure 3] FIG. 3 is a diagram showing an example of input and output of an analysis system according to a reference example. [Figure 4] FIG. 4 is a flowchart showing the procedure for generating n networks in the reference example. [Figure 5] FIG. 5(A) is a graph showing the distribution of the number of people who interacted, and FIG. 5(B) is an array showing the distribution of the number of people who interacted. [Figure 6] FIG. 6 is an explanatory diagram of branch allocation in the reference example. [Figure 7] FIG. 7 is a diagram illustrating an example of input and output of the analysis system according to the embodiment. [Figure 8] FIG. 8 is a flowchart showing a procedure for generating n networks in the embodiment. [Figure 9]FIG. 9 is a diagram illustrating analysis target data according to the embodiment. [Figure 10] FIG. 10 is an explanatory diagram of branch allocation in the embodiment. [Figure 11] FIG. 11 is a diagram showing the generated network. [Figure 12] FIG. 12 is a diagram illustrating an analysis example. [Figure 13] FIG. 13(A) is an explanatory diagram of the number of triangles in the entire network, and FIG. 13(B) is an explanatory diagram of triangles. [Figure 14] FIG. 14 is a diagram showing an example of an analysis result based on n networks. [Figure 15] FIG. 15 is an explanatory diagram of the probability of belonging to a triangle and the expected value of the number of triangles to which a vertex belongs. [Figure 16] FIG. 16 shows a visualized diagram of the network (display network). DETAILED DESCRIPTION OF THE INVENTION
[0011] 1. Overview of the Analysis System, Computer Implementation Method, and Computer Program
[0012] (1) An analysis system according to an embodiment may include a processor that performs operations including acquiring analysis target data, which is data relating to multiple elements, randomly generating multiple networks, and performing analysis based on the multiple networks.
[0013] The analysis target data may include the number of other elements associated with each of the plurality of elements. The number of other elements may indicate the number of other elements related to each of the plurality of elements. Each of the plurality of randomly generated networks may have a vertex in one-to-one correspondence with the plurality of elements. Each vertex may have a number of branches corresponding to the number of other elements associated with the element corresponding to each vertex. Because each of the plurality of randomly generated networks has a vertex in one-to-one correspondence with the plurality of elements, the vertices can be distinguished and analyzed. Furthermore, because each vertex has a number of branches corresponding to the number of other elements associated with the element corresponding to each vertex, a network appropriate for the number of other elements is generated.
[0014] (2) Randomly generating the plurality of networks may include allocating to each vertex in each of the plurality of networks generated a number of branches corresponding to the number of other elements associated with the element corresponding to each vertex, and randomly connecting the branches allocated to each vertex.
[0015] (3) The analysis target data may further include an element identifier for identifying each of the plurality of elements. Each element identifier may be associated with a respective number of other elements. Randomly generating the plurality of networks may include associating the element identifier with each of vertices, the number of which corresponds to the number of the plurality of elements, in each of the plurality of networks to be generated, allocating to each vertex a number of branches corresponding to the number of other elements associated with the element identifier associated with the vertex, and randomly connecting the branches allocated to each vertex.
[0016] (4) The analysis target data may further include characteristic data, which may be data indicating characteristics of each element other than the number of other elements.
[0017] (5) Performing an analysis based on the plurality of networks may include analyzing the plurality of networks using the feature data.
[0018] (6) Analyzing the plurality of networks using the feature data may include calculating, using the feature data, feature quantities related to vertices based on the plurality of networks.
[0019] (7) Analyzing the plurality of networks using the feature data may include calculating, based on the plurality of networks, feature amounts of each of a plurality of vertices using the feature data.
[0020] (8) Analyzing the plurality of networks using the feature data may include determining at least one of a first group of vertices that satisfy a first condition that can be determined from the feature data and a second group of vertices that do not satisfy the first condition, and calculating feature quantities for at least one of the first group and the second group based on the plurality of networks.
[0021] (9) Analyzing the plurality of networks may include calculating a characteristic amount of each of a plurality of vertices in the plurality of networks, and determining a group of vertices that satisfy a condition that can be determined from the characteristic amount.
[0022] (10) Analyzing the plurality of networks using the feature data may include calculating feature values for each of the plurality of vertices in the plurality of networks, determining a group of vertices that satisfy conditions that can be determined from the feature values, and analyzing the vertices included in the group using the feature data.
[0023] (11) Performing the analysis based on the plurality of networks may include calculating feature quantities related to branches based on the plurality of networks.
[0024] (12) Performing an analysis based on the plurality of networks may include calculating a probability of existence of an edge based on the plurality of networks.
[0025] (13) Performing analysis based on the plurality of networks may include generating a display network that indicates the existence probability of the edge as a connection weight between vertices.
[0026] (14) A method according to an embodiment may be a computer-implemented method executed by a computer. The computer-implemented method according to an embodiment may include acquiring analysis target data, which is data related to a plurality of elements, randomly generating a plurality of networks, and performing analysis based on the plurality of networks. The analysis target data may include the number of other elements associated with each of the plurality of elements. The number of other elements may indicate the number of other elements associated with each of the plurality of elements. Each of the plurality of randomly generated networks may have vertices in one-to-one correspondence with the plurality of elements. Each vertex may have a number of branches corresponding to the number of other elements associated with the element corresponding to the vertex.
[0027] (15) A computer program according to an embodiment causes a computer to execute an operation. The operation may include acquiring analysis target data, which is data relating to a plurality of elements, randomly generating a plurality of networks, and performing analysis based on the plurality of networks. The analysis target data may include the number of other elements associated with each of the plurality of elements. The number of other elements may indicate the number of other elements associated with each of the plurality of elements. Each of the plurality of randomly generated networks may have vertices in one-to-one correspondence with the plurality of elements. Each vertex may have a number of branches corresponding to the number of other elements associated with the element corresponding to the vertex.
[0028] 2. Examples of analysis systems, computer implementation methods, and computer programs
[0029] Examples of analysis systems, methods, and computer programs are described in more detail below with reference to the drawings.
[0030] FIG. 1 shows an analysis system 10 according to an embodiment. The analysis system 10 performs analysis of analysis target data. The analysis system 10 randomly generates multiple networks using analysis target data, which is data on each of multiple elements, and performs analysis based on the multiple networks. The analysis target data can be analyzed by analysis based on the multiple networks.
[0031] Here, with regard to the data to be analyzed, an "element" can refer to anything that can interact with other elements in real or virtual space, such as a natural person (individual), a corporation, an organization consisting of multiple natural persons or corporations, or an avatar in a virtual space such as the Metaverse. As an example, "multiple elements" are multiple individuals (members) belonging to a group (community). In the following description, as an example, an "element" is described as an "individual."
[0032] The relationships between "multiple elements" (for example, the social connections between multiple individuals) can be represented by a network. Here, "network" is synonymous with "graph" in graph theory, and is represented by vertices and edges connecting the vertices. For example, in a network that represents the connections between multiple individuals (elements), the vertices represent the individuals (elements), and the edges represent the connections between the individuals (elements). It can also be said that an "element" can be represented by the vertices of a network.
[0033] Networks representing groups of multiple individuals are useful for assessing the social connections of each individual. Assessing social connections can be useful, for example, for evaluating the degree to which individuals have achieved well-being. Well-being refers to being in good physical, mental, and social condition, and being healthy in a broad sense. Efforts are beginning to be made in various places to realize a society in which people can experience well-being.
[0034] To evaluate social connections, it is sufficient to generate a network that represents the connections and analyze the characteristics of that network. However, in order to represent social connections as a network, it is necessary to identify with whom each individual interacts. Furthermore, information about who an individual interacts with is personal information, so it can be difficult to collect. Furthermore, accurate investigations are often difficult due to the existence of people with the same name.
[0035] On the other hand, survey data on the number of people with whom an individual interacts, such as the number of friends, is relatively easy to collect. Therefore, in the embodiment, the analysis target data for a plurality of individuals is not data on the specific people with whom each individual interacts, but data simply indicating the number of interactions with each individual. That is, the analysis target data for the embodiment includes, as an example, number-of-interactions data indicating the number of other individuals related to each of the plurality of individuals.
[0036] The "interaction number data" here is an example of the "number of other elements." The "number of other elements" indicates the number of other elements related to each element, and for example, if the element is an individual, it indicates the number of people with whom each individual has interactions. "Related to each element" means that there is some kind of relationship, such as interaction with each element. As mentioned above, an "element" may be a company or the like other than an individual. For example, if the element includes a company, the "number of other elements" may be the number of other companies or individuals related to a certain company.
[0037] A network representing the connections between multiple individuals cannot be directly generated from analysis target data that does not include information about who the individuals interact with. If a network cannot be generated, the network's features cannot be determined. Therefore, the analysis system 10 according to the embodiment randomly generates multiple networks that are consistent with the data on the number of people interacting with the individual, and performs analysis using the multiple networks generated. This point will be described in detail later.
[0038] As shown in FIG. 1, the analysis system 10 may be configured with one or more computers. FIG. 1 shows an example in which the analysis system 10 is configured with one computer. The computer is, for example, a server computer or a personal computer. As an example, the analysis system 10 may be configured with a web server connected to a user's client terminal via a network. In this case, the analysis service according to the embodiment may be provided to the user as a web service.
[0039] The computer that constitutes the analysis system 10 includes a processor 20 and a storage device 30. The processor 20 is, for example, a CPU. The processor 20 is connected to the storage device 30.
[0040] The storage device 30 includes, for example, a primary storage device and a secondary storage device. The primary storage device is, for example, RAM. The secondary storage device is, for example, a hard disk drive (HDD) or a solid state drive (SSD). The storage device 30 includes a computer program 31 executed by the processor 20. The processor 20 reads and executes the computer program 31 stored in the storage device 30. The computer program 31 has program code that indicates instructions for causing a computer to execute information processing for analysis 21 of the analysis target data.
[0041] Data can be stored in the storage device 30. The data stored in the storage device 30 includes, for example, analysis target data 33, setting data 35, and analysis result data 37. The analysis target data 33 is data that is the subject of analysis 21 by the analysis system 10. The analysis target data 33 is, for example, acquired by the analysis system 10 from outside and stored in the storage device 30. The setting data 35 is data set for the analysis 21, such as the conditions for the analysis 21. The setting data 35 may be set by a user of the analysis system 10, or may be initially set. The analysis result data 37 is data that indicates the results of the analysis 21. The analysis result data 37 is generated by the analysis system 10 analyzing the analysis target data, and is stored in the storage device 30.
[0042] The analysis system 10 may include an input device 40. The input device 40 is a device for inputting data from the outside into the analysis system 10. The input device 40 may be used, for example, to input analysis target data 33 into the analysis system 10. The input device 40 may be a device that inputs data through user operation, such as a keyboard or mouse, or may be a data communication device that acquires data from the outside via communication.
[0043] The analysis system 10 may include an output device 50. The output device 50 is a device that outputs data. The output device 50 may be used, for example, to output analysis result data 37. The output device 50 may be, for example, a data display device, such as a display device, or may be a data communication device that transmits data to the outside.
[0044] 2 shows the procedure for performing analysis 21 by analysis system 10. In step S21, analysis system 10 acquires data. The acquired data is stored in storage device 30. The data acquired in step S21 includes, for example, analysis target data and setting data. If the setting data has been stored in analysis system 10 in advance, there is no need to acquire the setting data in step S21.
[0045] In step S22, the analysis system 10 determines the number n of networks to be randomly generated. Although each of the plurality of randomly generated networks does not accurately represent the connection between individuals, if the number n of randomly generated networks is sufficiently large, the expected values such as the characteristics of the network can be accurately obtained. The required number n of networks can be determined according to the accuracy of the analysis results. The number n of networks or the accuracy can be specified by the user as setting data. As an example, the number n of networks can be determined using Hoeffding's inequality according to what percentage of probability the error of the feature amount obtained as the analysis result is to be within what percentage.
[0046] For example, n can be obtained by satisfying Hoeffding's inequality: 2exp(-2nptol^2) < etol. Here, etol is the allowable error for the analysis result, and ptol is the allowable probability. When using the n networks determined in this way, an analysis result with an accuracy of a relative error etol can be obtained with a probability of (1 - ptol) or more. Here, the relative error indicates the relative error of the analysis result with respect to (b - a) when the maximum value of the analysis result (network feature amount) is b and the minimum value is a.
[0047] For example, when it is desired to make the error 1% or less with a probability of 95%, in the above Hoeffding's inequality, ptol = 0.05 and etol = 0.01 may be set. In this case, n can be approximately 18500.
[0048] In step S23, the analysis system 10 randomly generates n networks. As an algorithm for randomly generating networks, the configuration model can be used. If there is any assumption about the distribution of the network, an appropriate random graph model corresponding thereto can be used. The details of the algorithm for randomly generating networks will be described later.
[0049] In step S24, analysis system 10 performs analysis based on the generated n networks. Analysis system 10 stores the analysis result data obtained by the analysis in step S24 in storage device 30. Details of the analysis based on the n networks will be described later.
[0050] In step S25, the analysis system 10 outputs the analysis results of step S24. The output of the analysis results may be a display of the analysis result data on a computer screen, or may be a transmission of the analysis result data to an external device.
[0051] 3 to 6 show analysis 21 according to a reference example. To facilitate understanding of analysis 21 according to the embodiment, analysis 21 according to the reference example will be described first. The reference example conforms to the method described in Non-Patent Document 1. That is, in the reference example, the results of a survey on the number of interactions of each member of a community are treated as a degree distribution of the vertices of a network, a network conforming to such a degree distribution is randomly generated, and network features are estimated.
[0052] 3, in the reference example, the total number of members (number of elements) N and the distribution of the number of people interacting d0[k] are input as analysis target data to the analysis system 10. In addition, the allowable error etol, the allowable probability ptol, and the type of network feature specified by the user are input as setting data to the analysis system 10.
[0053] To analyze the data to be analyzed, the analysis system 10 determines n, randomly generates n networks based on the data to be analyzed, performs analysis based on the n networks, and outputs, as the analysis results, expected values of the types of network features specified by the user (see Figure 2).
[0054] Fig. 4 shows a method for generating n networks in a reference example. In the procedure of Fig. 4, the analysis system 10 generates n networks by repeatedly executing the network generation of step S40 n times. After generating n networks, the analysis system 10 ends the processing of Fig. 4 (step S44).
[0055] One network is randomly generated by executing step S40 once. In step S40, the analysis system 10 first generates N vertices corresponding to the total number of members N (step S41). Next, the analysis system 10 assigns edges to each of the N vertices using the interaction number distribution d0[k] (step S42). Then, the analysis system 10 randomly connects the edges assigned to each vertex (step S43). One network is generated by the above process. Then, by repeating the above process n times, n networks are generated.
[0056] Steps S42 and S43 will be described in more detail below. FIG. 5 shows an example of the interaction number distribution d0 used in step S42. This interaction number distribution d0 is created, for example, from a survey response regarding the number of interactions for each member. In this case, the interaction number distribution is represented by the number of respondents for each interaction number (the number of people with whom each member interacts within the community). FIG. 5(A) is a graph showing the interaction number distribution d0, with the horizontal axis representing the interaction number k and the vertical axis representing the number of respondents d0[k]. FIG. 5(B) shows the interaction number distribution of FIG. 5(A) in array data format d0[k]. As shown in FIG. 5(B), the number of people who answered that the number of interactions is 0 (k=0) is d0[0]=39, the number of people who answered that the number of interactions is 1 (k=1) is d0[1]=85, and the number of people who answered that the number of interactions is 2 (k=2) is d0[2]=107.
[0057] In the reference example, the interaction number distribution d0 is regarded as the degree distribution of the network. The degree distribution of a network indicates the distribution of the degrees of each vertex included in the network. Here, the "degree of a vertex" refers to the number of edges (edges) connected to a vertex. When the interaction number distribution d0 is regarded as the degree distribution of the network, the graph in FIG. 5(A) is regarded as having the degree k on the horizontal axis and the number of vertices on the vertical axis. Furthermore, the array in FIG. 5(B) indicates the number of vertices d0[k] for each degree k. For example, the array in FIG. 5(B) is regarded as a degree sequence d0[k] indicating that there are 39 vertices with degree k=0, 85 vertices with degree k=1, 107 vertices with degree k=2, 84 vertices with degree k=3, 53 vertices with degree k=4, and 29 vertices with degree k=5.
[0058] In step S42 of FIG. 4, the analysis system 10 assigns edges to each of the N vertices corresponding to the total number of members N, using the degree sequence d0[k] of FIG. 5(B). FIG. 6 shows how edges are assigned to the N vertices. As shown in FIG. 6, the analysis system 10 assigns no edges (k=0 edges) to d0[0]=39 vertices out of the N vertices. The analysis system 10 assigns k=1 edges to d0[1]=85 vertices out of the N vertices. Similarly, the analysis system 10 assigns k=2 edges to d0[2]=107 vertices, k=3 edges to d0[3]=84 vertices, k=4 edges to d0[4]=53 vertices, and k=5 edges to d0[5]=29 vertices out of the N vertices. In this way, edges are assigned to all N vertices according to the degree sequence d0[k]. At the time of step S42, one end of each edge is connected to a vertex, but the other end is not connected to a vertex, and the connection destination is undetermined.
[0059] In step S43, the analysis system 10 randomly connects the edges assigned to each vertex, generating a network that is consistent with the degree distribution d0[k] by connecting the edges.
[0060] In order to connect edges randomly, in step S43, the analysis system 10 repeats the following processes (1) to (3) while the following v and w can be selected. When the connection destinations of all edges have been determined, v and w can no longer be selected, and the following processes end. (1) Randomly select two vertices v and w from among N vertices. (2) If the two vertices v and w include a vertex to which all edge destinations have been determined, reselect the two vertices v and w. (3) From each of the vertices v and w, select one branch that has no connection destination and connect those branches.
[0061] In the above (3), the edge assigned to vertex v and the edge assigned to vertex w are connected to form a single edge connecting vertex v and vertex w.
[0062] In the reference example, edges are simply assigned according to the degree distribution d0[k], and vertices with the same degree are not distinguished. For example, in Figure 6, the 85 vertices with degree k=1 are not distinguished from one another. Similarly, the 107 vertices with degree k=2 are not distinguished from one another.
[0063] Therefore, when using n networks generated by the reference examples (Figures 3 to 6), it is possible to analyze global characteristics of the entire network (characteristics of the entire group represented by the network), such as the clustering coefficient or average path length, but it is not possible to distinguish between vertices with the same degree. As a result, it is not possible to analyze, for example, individuals belonging to a group represented by a network.
[0064] In contrast, by using n networks generated using the methods shown in Figures 7 to 10, it is possible to distinguish between each of the vertices (individuals) included in the network, even if they have the same degree, and analyze the network's characteristics. Analyzable characteristics include, for example, the probability that each individual is included in a cluster, the number of clusters that each individual is included in, and the importance of each individual in the network. Furthermore, because it is possible to distinguish between the vertices, it is also possible to distinguish between the branches that connect the vertices. Therefore, it is possible to analyze the network's characteristics by distinguishing between the branches.
[0065] Furthermore, in the case of the methods shown in Figures 7 to 10, the data to be analyzed contains feature data for individuals (elements), so it is possible to calculate the average value of the feature quantity for each vertex to be analyzed for each value indicating the feature (qualities) of the individual to be investigated, such as gender, age, personality, etc., to determine whether or not it has an impact on the state of interpersonal interactions. This allows the conditional expected value for the feature quantity to be found depending on gender, age, personality, etc. As a result, it becomes possible to analyze how gender, age, personality, etc. affect the state of interpersonal interactions.
[0066] 7 to 10 are explanatory diagrams of an analysis 21 according to an embodiment. In the example shown in FIG. 7, the total number of members (number of elements) N, a feature table m[j, h], and data on the number of people interacting d1[j] are input to the analysis system as analysis data. In addition, the allowable error etol, the allowable probability ptol, the significance level alpha of the test, the type of network feature specified by the user, a first condition C, a second condition D, and a third condition E are input to the analysis system 10 as setting data. Note that it is not necessary to input all of the setting data shown in FIG. 7; only the setting data necessary for the analysis to be performed is sufficient.
[0067] To analyze the data to be analyzed, the analysis system 10 determines n, randomly generates n networks based on the data to be analyzed, performs analysis based on the n networks, and outputs the analysis results, which may include, for example, expected values of network features, inverse analysis results, statistical test results, and visualized diagrams of the networks.
[0068] Fig. 8 shows a method for generating n networks in an embodiment. In the procedure of Fig. 8, the analysis system 10 generates n networks by repeatedly executing the network generation of step S80 n times. After generating n networks, the analysis system 10 ends the processing of Fig. 8 (step S85).
[0069] One network is randomly generated by executing step S80 once. In step S80, the analysis system 10 first generates N vertices corresponding to the total number of members N (step S81). Next, the analysis system 10 assigns edges to each of the N vertices j according to the number of people interacting data d1[j] (step S82). The analysis system 10 also associates a feature [j, h] with each vertex j. Then, the analysis system 10 randomly connects the edges assigned to each vertex (step S85). One network is generated by the above process. Then, by repeating the above process n times, n networks are generated.
[0070] Steps S82 to S84 will be described in more detail below. FIG. 9 shows an example of analysis target data including number of contacts data d1[j] and feature data m[j,h]. This analysis target data is created, for example, from questionnaire responses to each member. The questionnaire may include various questionnaire items related to the number of contacts of each member, as well as each member's gender, age, personality, habits, and preferences. These questionnaire items indicate the characteristics of each member (individual). In FIG. 9, h is an identifier for a member's feature other than the number of contacts; for example, h=1 indicates "gender" and h=2 indicates "age." Among the data indicating gender, "1" and "2" indicate female and male, respectively.
[0071] The analysis target data shown in FIG. 9 has a member ID (element ID; element identifier) j for identifying each member (element). In the analysis target data shown in FIG. 9, the member ID (element identifier) is associated with the number of interactions (number of other elements) of that member. For example, a member (element) with member ID (element identifier) j=0 is associated with the number of interactions (number of other elements) d1[0]=3. In FIG. 9, the pair of member ID (element identifier j) and the number of interactions (number of other elements) constitutes the number-of-interactions data d1[j].
[0072] In addition, in the analysis target data shown in FIG. 9, the member ID (element identifier) is associated with the member's characteristics (h=1 to 4). For example, a member with member ID j=0 is associated with gender = 1 (gender = 1 indicates, for example, female) and age = 64. In FIG. 9, a pair of member ID (element identifier) j and a characteristic (h=1 to 4) other than the number of contacts constitutes characteristic data m[j,h]. For example, m[0,1] indicates that the member with member ID j=0 is female.
[0073] In step S82 of FIG. 8, the analysis system 10 assigns N vertices, corresponding to the total number of members, to the N members in a one-to-one correspondence, and then assigns edges to each vertex in accordance with the number-of-interactions data d1[j]. For example, the member ID (element identifier) j is also used as a vertex identifier, and the number of edges assigned to each vertex j corresponds to the number of interactions (number of other elements) indicated by the number-of-interactions data d1[j] (step S82). Furthermore, by using the member ID (element identifier) j as a vertex identifier, a feature m[j,h] is associated with each vertex j (step S83). By associating a feature m[j,h] with each vertex j, it becomes possible to use the feature m[j,h] when analyzing multiple networks.
[0074] Figure 10 shows an example in which branches are assigned to six vertices V1, V2, V3, V4, V5, and V6 based on the number of people interacting data d1[j]. In Figure 10, the first vertex V1 is associated with member ID j=0. Therefore, the first vertex V1 indicates the member with j=0, and the number of people interacting d1[0]=3 and the feature m[0,h] are associated with the first vertex V1. Three branches are assigned to the first vertex V1 according to the number of people interacting d1[0].
[0075] Similarly, the second vertex V2 is associated with member ID j=1. The second vertex V2 is associated with the number of interactions d1[1]=0 and the feature m[1,h], and is assigned 0 branches. The third vertex V3 is associated with member ID j=2. The third vertex V3 is associated with the number of interactions d1[2]=0 and the feature m[2,h], and is assigned 0 branches. The fourth vertex V4 is associated with member ID j=3. The fourth vertex V4 is associated with the number of interactions d1[3]=2 and the feature m[3,h], and is assigned 2 branches. The fifth vertex V5 is associated with member ID j=4. The fifth vertex V5 is associated with the number of interactions d1[4]=5 and the feature m[4,h], and is assigned 5 branches. The sixth vertex V6 is associated with member ID j=5. The sixth vertex V6 is associated with the number of contacts d1[5]=3 and the feature m[5,h], and three branches are allocated to it.
[0076] In this way, in step S82, branches are allocated to all N vertices, the number of which corresponds to the number of members who interact with each vertex. At the time of step S82, one end of each branch is connected to a vertex, but the other end is not connected to a vertex, and the connection destination is undetermined.
[0077] In step S84, the analysis system 10 randomly connects the edges assigned to each vertex. By connecting the edges randomly, a network that is consistent with the number of people interacting data d1[j] is randomly generated.
[0078] In step S84, the procedure for randomly connecting edges is the same as in step S43. That is, in step S84, the analysis system 10 repeats the following processes (1) to (3) while the following v and w can be selected. When the connection destinations of all edges have been determined, v and w can no longer be selected, and the following processes end. (1) Randomly select two vertices v and w from among N vertices. (2) If the two vertices v and w include a vertex to which all edge destinations have been determined, reselect the two vertices v and w. (3) From each of the vertices v and w, select one branch that has no connection destination and connect those branches.
[0079] By randomly connecting the edges, a network such as that shown in FIG. 11 is generated.
[0080] Each of the n randomly generated networks has N vertices j, which correspond one-to-one to N members (elements), and each vertex j has a number of branches corresponding to the number of people d1[j] (the number of other elements) associated with the member (element) j corresponding to each vertex j. The number of vertices (N) is the same across all n networks, and the degree of each vertex j is also the same, but the connections between vertices by the branches are randomly determined. Therefore, the connections between vertices via the branches vary from network to network. Although the branches indicate the "connections" between members, the branches in each of the n networks do not necessarily reflect the actual "connections" between members. However, by analyzing a sufficiently large number of networks, n, it becomes possible to reliably estimate the "connections."
[0081] In addition, in the network generated by the procedure in Figure 8, member (element) identifiers are associated with each vertex, so even if there are multiple vertices with the same degree, they can be distinguished. Therefore, by analyzing n networks, it is possible to analyze the members corresponding to the vertices.
[0082] Furthermore, because individual vertices can be distinguished, the edges connecting the vertices can also be distinguished individually. It is also possible to identify whether an edge exists between any two vertices. Therefore, by analyzing n networks, it becomes possible to analyze (and visualize) the "connections between members (individuals)" that correspond to the edges.
[0083] FIG. 12 shows an example of analysis (step S24) that can be performed by the analysis system 10 according to the embodiment. As shown in FIG. 12, the analysis system 10 can use n networks to determine various types of feature quantities. The type of feature quantity to be determined can be set by user input, for example, as "designation of network feature quantity" (see FIG. 7). The analysis system 10 can execute processing to determine the type of feature quantity designated by "designation of network feature quantity" from the feature quantities shown in FIG. 12.
[0084] Step S121 of the process shown in FIG. 12 is a process for calculating the feature of the entire network. The feature of the entire network is a global feature (a feature of a group represented by a network) seen from the perspective of the entire network, such as a cluster coefficient or an average path length. The analysis system 10 calculates the feature of the entire network as the average value of the feature of n networks. The average value of the feature of n networks indicates the expected value of the feature. If n is sufficiently large, the expected value of the feature will be sufficiently reliable. The expected value is calculated with a probability of (1-ptol) or greater and an accuracy of a relative error of etol%. Note that the feature of the entire network can also be calculated in the reference examples shown in FIGS. 3 to 6.
[0085] Figure 13 shows the "number of triangles" as another example of a feature of the entire network (global network feature). Here, a "triangle" refers to a connection relationship in which three vertices are connected to each other by edges, as shown in Figure 13(B). The "number of triangles" refers to the number of triangles that exist in an entire network. In the case of the network in Figure 13(A), the number of triangles is two. Because a triangle represents a mutual helping structure involving three parties, the number of triangles indicates the number of mutual helping structures. Therefore, the number of triangles is useful for understanding the number of mutual helping structures in the group represented by the network.
[0086] To calculate the expected number of triangles, we calculate the number of triangles for each of n networks and take the average of the number of triangles as the expected value. When n is sufficiently large, the expected number of triangles reliably indicates the expected number of mutual helping structures in the group.
[0087] In Fig. 12, the process of step S121 calculates features (global network features) that can be obtained without distinguishing between individual vertices, while the processes of steps S122 and S127 calculate features that can be obtained by distinguishing between individual vertices. In step S122, features related to vertices are calculated, and in step S127, features related to edges are calculated. The value obtained in step S121 is called the "expected value related to the feature," and the value obtained in step S122 or step S127 is called the "conditional expected value related to the feature."
[0088] In FIG. 12, four types of processing steps S123, S124, S125, and S126 are illustrated as processing steps for calculating feature amounts related to vertices.
[0089] The process of step S123 in FIG. 12 is a process for calculating feature values for each individual corresponding to a vertex. FIG. 14 shows an example of feature values for each individual. In FIG. 14, for each member (individual), the probability of belonging to a triangle (see FIG. 13(B)) and the expected value of the number of triangles to which each member belongs are calculated based on n networks. The "probability of belonging to a triangle" indicates the probability that each member is included in a triangle in n networks. The "expected value of the number of triangles to which each member belongs" is the average number of triangles to which each member belongs in n networks. The analysis results obtained in this manner are stored in the storage device 30 in association with the member ID (element ID), as shown in FIG. 14. Therefore, the analysis results for each member are associated with the feature data for each member. As a result, the analysis system 10 can analyze the relationship between each member's features, such as age or gender, and the nature of their interactions (e.g., the number of triangles), based on the feature data. In this way, the process of step S123 may not simply distinguish each vertex and calculate feature values, but may also analyze the feature values of each vertex using the feature data associated with each vertex.
[0090] The process of step S124 in FIG. 12 is a process of calculating a feature value for each group of vertices using feature data associated with each vertex. In step S124, for example, at least one of a first group of vertices that satisfy a first condition C that can be determined from the feature data m[j,h] and a second group of vertices that do not satisfy the first condition C is obtained, and a feature value for at least one of the first group and the second group is calculated based on n networks. The first condition C that can be determined from the feature data m[j,h] is, for example, a condition related to gender or age, specifically, "being female," "being male," "being 65 years or older," etc. Note that the first condition does not need to be based on only one piece of feature data m[j,h], but may be a combination of multiple piece of feature data m[j,h] (for example, a combination of AND or OR).
[0091] FIG. 15 shows the calculation results of feature quantities when the first condition C is "65 years old or older" in one network shown in FIG. 15. Here, the first group that satisfies the first condition C is a set of vertices corresponding to individuals whose age is 65 years old or older. The second group that does not satisfy the first condition C is a set of vertices corresponding to individuals whose age is younger than 65 years old. Note that in FIG. 15, the numbers in circles that indicate the vertices of the network indicate the ages of the individuals corresponding to the vertices.
[0092] Based on the first condition C, the analysis system 10 refers to the feature data m[j,h] of the individual corresponding to each vertex, and obtains a first group N1 which is a first set of vertices that satisfy the first condition C, and a second group N2 which is a second set of other vertices. Here, the number of vertices included in the first group N1 is defined as n1, and the number of vertices included in the second group N2 is defined as n2.
[0093] The analysis system 10 calculates the feature quantity of the type specified by the user for each vertex included in the first group N1 based on the n networks. The analysis system 10 also calculates the feature quantity of the type specified by the user for each vertex included in the second group N2 based on the n networks. Here, the feature quantities specified by the user are the "probability of belonging to a triangle" and the "expected value of the number of triangles belonging to the triangle."
[0094] For example, the analysis system 10 calculates the feature values for each of the n networks for the vertices included in the first group N1, and calculates the average value c1. Then, c1 is updated to c1 / (n1+n2) / n1. Furthermore, the analysis system 10 calculates the feature values for each of the n networks for the vertices included in the second group N2, and calculates the average value c2. Then, c2 is updated to c2 / (n1+n2) / n2.
[0095] For example, in the network shown in FIG. 15, the probability that a vertex (individual) aged 65 or older belongs to a triangle is 4 / 5, and the probability that a vertex (individual) under 65 belongs to a triangle is 1 / 3. The expected value of the number of triangles containing vertices (individuals) aged 65 or older is 4 / 5, and the expected value of the number of triangles containing vertices (individuals) under 65 is 2 / 3. By calculating these feature amounts for each of the n networks and averaging them, it is possible to determine the network feature amounts for each group according to the first condition C. For example, it is possible to analyze the relationship between differences in age or gender and the number of triangles (the number of mutually helping structures).
[0096] The process of step S125 in Fig. 12 is a process of performing an inverse analysis using the feature data associated with each vertex. Here, the process of step S124 is a type of analysis in which "people who have the feature of ... have the network feature of ...." For example, the process of step S124 is a process for obtaining an analysis result that people aged 65 or older (people who have the personal feature of being 65 or older) have a network feature in which the number of triangles (the number of mutually helping structures) is large.
[0097] In contrast to this, the processing in step S125 is an analysis (reverse analysis) of the type "people who have a network characteristic of ~ tend to have a personal characteristic of ~." For example, this is processing to obtain an analysis result that people who have a network characteristic of a large number of triangles (number of mutually helping structures) tend to be 65 years old or older (people who have the personal characteristic of being 65 years old or older).
[0098] In step S125, the analysis system 10 calculates the network feature of each of the multiple vertices j in each of the n networks, and finds a group of vertices that satisfies a second condition D that can be determined from the network feature. This results in a group that satisfies the second condition D. Also in step S125, the analysis system 10 analyzes the vertices included in the group that satisfies the second condition D using the feature data m[j,h].
[0099] In step S125, calculation is performed using, for example, a second condition D that can be determined from the network feature determined at each vertex, and a third condition E that can be determined from the feature data m[j,h]. The second condition D is, for example, a condition related to betweenness centrality, page rank, number of triangles, etc. The third condition E is, for example, a condition related to gender or age, specifically, "is female," "is male," "is 65 years old or older," etc. Note that the third condition E does not need to be based on only one piece of feature data m[j,h], but may be a combination of multiple piece of feature data m[j,h] (for example, a combination of AND or OR).
[0100] More specifically, the analysis system 10 determines a group N3 of vertices that satisfy the second condition D. Here, the number of vertices included in group N3 is n3. Furthermore, the analysis system 10 determines the number n4 of vertices in group N3 that satisfy the third condition E. The analysis system 10 also determines pe = n4 / n3. The analysis system 10 determines the average value of pe in n networks. The average value of pe indicates the probability that vertices that satisfy condition D also satisfy condition E. The average value of pe indicates, for example, the probability that a vertex (individual) with a predetermined number of triangles or more (condition D) is 65 years old or older (condition E).
[0101] The process of step S126 in Fig. 12 is a process of performing a statistical test. In the statistical test, as an example, a test result is obtained with the third condition E as the null hypothesis. In the statistical test, as an example, the result of the inverse analysis of step S125 (probability pe: the probability that a vertex that satisfies condition D satisfies condition E) is used. If the analysis system 10 determines that probability pe < the significance level alpha of the test is satisfied, it can reject the null hypothesis.
[0102] 12, the process of step S128 is illustrated as an example of the process of calculating the feature amount related to the edge. The process of step S128 is a process of calculating the existence probability of an edge in n networks.
[0103] In step S128, the analysis system 10 calculates p ij Here, i and j respectively represent the vertices included in the network. ij indicates whether or not there is an edge between vertex i and vertex j. Here, as an example, if there is an edge between vertex i and vertex j, the analysis system 10 determines whether p ij If there is no edge between vertex i and vertex j, then p ij is set to "0".
[0104] The analysis system 10 calculates p in n networks. ij The average value of is calculated as the existence probability of each edge. The existence probability of an edge indicates the probability of the connection between the individuals corresponding to the two vertices connected by the edge.
[0105] The analysis system 10 can generate a display network as shown in FIG. 16 based on the existence probability of a branch. The display network shown in FIG. 16 is a network that shows the existence probability of a branch as a connection weight between vertices. The connection weight indicates an estimated value of the degree of connection between individuals. Therefore, the display network visualizes the estimated results of the connection in the group to which the individual belongs. In FIG. 16, the connection weight is indicated by the thickness of the branch, but the thickness of the branch may be common and a numerical value indicating the connection weight (existence probability) may be displayed near the branch. Furthermore, the connection weight may be indicated by the thickness of the branch and also displayed as a numerical value.
[0106] The present invention is not limited to the above-described embodiment, and various modifications are possible. [Explanation of symbols]
[0107] 10: Analysis system 20: Processor 21: Analysis 30: Storage device 31: Computer Program 33: Data to be analyzed 35: Setting data 37: Analysis result data 40: Input device 50: Output device
Claims
1. Acquire data to be analyzed, which is data relating to multiple elements, Randomly generate multiple networks, performing an analysis based on the plurality of networks; a processor that performs operations including: the analysis target data includes the number of other elements associated with each of the plurality of elements, the number of other elements indicates the number of other elements related to each of the plurality of elements, Each of the plurality of randomly generated networks has vertices that correspond one-to-one to the plurality of elements, and each vertex has branches the number of which corresponds to the number of other elements associated with the element corresponding to each vertex. Analysis system.
2. Randomly generating the plurality of networks includes: In each of the plurality of networks to be generated, a number of branches corresponding to the number of other elements associated with the element corresponding to each vertex is allocated to each vertex, and the branches allocated to each vertex are randomly connected. Including, The analysis system according to claim 1 .
3. the analysis target data further includes an element identifier for identifying each of the plurality of elements, and each element identifier is associated with the number of other elements; Randomly generating the plurality of networks includes: In each of the plurality of networks to be generated, Associating the element identifier with each of the vertices, the number of which corresponds to the number of the plurality of elements; Allocating to each vertex a number of branches corresponding to the number of other elements associated with the element identifier associated with the vertex; Randomly connect the edges assigned to each vertex. Including, The analysis system according to claim 1 .
4. The analysis target data further includes feature data, The feature data is data indicating features of each element other than the number of other elements. The analysis system according to any one of claims 1 to 3.
5. performing an analysis based on the plurality of networks includes analyzing the plurality of networks using the feature data; The analysis system according to claim 4 .
6. analyzing the plurality of networks using the feature data includes calculating, based on the plurality of networks, feature amounts related to vertices using the feature data; The analysis system according to claim 5 .
7. Analyzing the plurality of networks using the feature data includes calculating, based on the plurality of networks, feature amounts of each of a plurality of vertices using the feature data. The analysis system according to claim 5 .
8. Analyzing the plurality of networks using the feature data includes determining at least one of a first group of vertices that satisfy a first condition that can be determined from the feature data and a second group of vertices that do not satisfy the first condition, and calculating a feature amount related to at least one of the first group and the second group based on the plurality of networks. The analysis system according to claim 5 .
9. Analyzing the plurality of networks includes calculating a feature amount of each of a plurality of vertices in the plurality of networks, and determining a group of vertices that satisfy a condition that can be determined from the feature amount. The analysis system according to claim 1 .
10. Analyzing the plurality of networks using the feature data includes calculating feature amounts of each of a plurality of vertices in the plurality of networks, determining a group of vertices that satisfy a condition determinable from the feature amounts, and analyzing the vertices included in the group using the feature data. The analysis system according to claim 5 .
11. performing an analysis based on the plurality of networks includes calculating feature amounts related to branches based on the plurality of networks; The analysis system according to claim 1 .
12. performing an analysis based on the plurality of networks includes calculating a probability of existence of an edge based on the plurality of networks; The analysis system according to claim 1 .
13. performing an analysis based on the plurality of networks includes generating a display network that indicates the existence probabilities of the edges as connection weights between vertices; The analysis system according to claim 12.
14. 1. A computer-implemented method performed by a computer, comprising: Acquire data to be analyzed, which is data relating to multiple elements, Randomly generate multiple networks, performing an analysis based on the plurality of networks; This includes: the analysis target data includes the number of other elements associated with each of the plurality of elements, the number of other elements indicates the number of other elements related to each of the plurality of elements, Each of the plurality of randomly generated networks has vertices that correspond one-to-one to the plurality of elements, and each vertex has branches the number of which corresponds to the number of other elements associated with the element corresponding to each vertex. Computer-implemented methods.
15. A computer program that causes a computer to perform an operation, The operation is Acquire data to be analyzed, which is data relating to multiple elements, Randomly generate multiple networks, performing an analysis based on the plurality of networks; This includes: the analysis target data includes the number of other elements associated with each of the plurality of elements, the number of other elements indicates the number of other elements related to each of the plurality of elements, Each of the plurality of randomly generated networks has vertices that correspond one-to-one to the plurality of elements, and each vertex has branches the number of which corresponds to the number of other elements associated with the element corresponding to each vertex. Computer program.
Citation Information
Patent Citations
Business microscope system
JP2008210363A
JP2022