Information processing apparatus, information processing method, and program

By clustering surrounding meshes and constructing separate estimation models for each group using base mesh data, the method addresses the inefficiencies in existing technologies, enhancing the efficiency of attribute value estimation for multiple regions.

JP2025158214APending Publication Date: 2025-10-17KK TOSHIBA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024060542
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-04
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently constructing estimation models for attribute values of multiple regions due to excessive load when modeling relationships between base and surrounding areas, particularly in scenarios where the number of surrounding meshes exceeds that of base areas.

Method used

The approach involves clustering surrounding meshes into multiple groups and constructing an estimation model for each cluster using base mesh data, reducing the processing load and improving efficiency.

Benefits of technology

This method reduces the time and cost required for model construction by efficiently estimating attribute values, such as people flow, by grouping surrounding meshes and building separate models for each cluster.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025158214000001_ABST
    Figure 2025158214000001_ABST
Patent Text Reader

Abstract

To efficiently construct an estimation model for estimating an attribute of a region included in a plurality of regions.SOLUTION: An information processing apparatus includes a processing unit. The processing unit generates a plurality of clusters for classifying a plurality of second regions, using multiple pieces of second attribute data representing attributes of the second regions other than one or more first regions included in a plurality of regions. The processing unit constructs an estimation model configured to receive, as input, one or more pieces of first attribute data representing attributes of the one or more first regions, and output cluster attribute data representing the clusters, for multiple clusters, using learning data including the first attribute data for learning and the second attribute data of the second regions classified into clusters.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] An embodiment of the present invention relates to an information processing device, an information processing method, and a program. [Background technology]

[0002] A technology has been proposed that uses artificial intelligence (AI), machine learning, traffic engineering, etc. to model the relationship between attributes (e.g., people flow) between a base area among multiple mesh-like areas and areas surrounding the base area, and then uses this model to estimate attribute values.

[0003] To improve the accuracy of estimation, it is necessary to learn features for each of multiple meshes. For example, when estimating attribute values ​​of a surrounding area from attribute data representing the attributes of a base area, it is necessary to model each of multiple meshes corresponding to the surrounding area. Since the number of meshes corresponding to the surrounding area is greater than the number of base areas, the load when building the model may become excessive. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Patent No. 5576838 Summary of the Invention [Problem to be solved by the invention]

[0005] An object of the present invention is to provide an information processing device, an information processing method, and a program that can more efficiently construct an estimation model for estimating the attributes of one of a plurality of regions. [Means for solving the problem]

[0006] An information processing apparatus according to an embodiment includes a processing unit. The processing unit generates a plurality of clusters for classifying the plurality of second regions using a plurality of second attribute data representing attributes of each of a plurality of second regions other than one or more first regions included in the plurality of regions. The processing unit constructs, for each of the plurality of clusters, an estimation model that inputs one or more first attribute data representing the attributes of each of the one or more first regions and outputs cluster attribute data representing the attributes of the cluster, using training data including the first attribute data for training and the second attribute data of the second regions classified into the cluster. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 10 is a diagram showing an example of a mesh configuration. [Figure 2] FIG. 1 is a block diagram showing the configuration of an information processing apparatus according to a first embodiment. [Figure 3] FIG. 4 is a diagram showing an example of the data structure of surrounding people flow data. [Figure 4] FIG. 10 is a diagram showing an example of the data structure of base point people flow data. [Figure 5] 10 is a flowchart of a model learning process. [Figure 6] 10 is a flowchart of a cluster generation process. [Figure 7] FIG. 4 is a diagram showing an example of the data structure of cluster information. [Figure 8] 10 is a flowchart of a clustering process according to the first embodiment. [Figure 9] 10 is a flowchart of a model construction process. [Figure 10] FIG. 10 is a diagram showing an example of learning data. [Figure 11] 4 is a flowchart of an estimation process according to the first embodiment. [Figure 12] FIG. 10 is a block diagram of an information processing apparatus according to a second embodiment. [Figure 13] FIG. 4 is a diagram showing an example of the data structure of surrounding mesh information. [Figure 14] 10 is a flowchart of a clustering process according to the second embodiment. [Figure 15] FIG. 10 is a block diagram of an information processing apparatus according to a third embodiment. [Figure 16] FIG. 4 is a diagram showing an example of the data structure of base point mesh information. [Figure 17] A diagram showing an overview of model construction using GNN. [Figure 18] FIG. 10 is a block diagram of an information processing apparatus according to a fourth embodiment. [Figure 19] FIG. 2 is a diagram showing an example of the data structure of input data used in classification processing. [Figure 20] 10 is a flowchart of a classification process. [Figure 21] FIG. 1 is a hardware configuration diagram of an information processing apparatus according to first to fourth embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0008] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Preferred embodiments of an information processing apparatus according to the present invention will be described in detail below with reference to the accompanying drawings.

[0009] In the following embodiments, an example will be described in which people flow values ​​that represent the population within a mesh are estimated. A mesh refers to a section of a geographical space divided into a grid pattern. The people flow value is, for example, the population (number of people) or population density. The data to be estimated is not limited to people flow values, and may be any data that represents the attributes of each area (hereinafter referred to as attribute data). For example, the attributes may be traffic volume (the number of moving objects such as vehicles), network traffic, power consumption, weather information, etc.

[0010] The people flow value can be interpreted as amount data representing the amount of people, which is an example of an object. When the object is a mobile object other than a person, network traffic, or power consumption, the amount data can be interpreted as traffic volume, network traffic, or power consumption.

[0011] Furthermore, the area that serves as the unit for estimating attribute data is not limited to a mesh (a mesh-shaped area), and may be any other area. As an area other than a mesh, for example, the following areas can be used. Areas divided into roughly equal sizes based on latitude and longitude Areas bounded by city, town, village, and address boundaries

[0012] FIG. 1 is a diagram showing an example of a mesh configuration. FIG. 1 shows an example of a mesh in which geographic space is divided into 8 x 8 = 64 grids. As shown in FIG. 1, the multiple meshes include a base mesh and surrounding meshes. In the example of FIG. 1, four regions 10a to 10d correspond to the base meshes, and the other 60 regions correspond to the surrounding meshes. The base mesh corresponds to one or more regions (first regions) that serve as base points among the multiple meshes. The surrounding meshes correspond to multiple regions (second regions) other than the base mesh among the multiple meshes.

[0013] In the following embodiment, the people flow values ​​of 60 surrounding meshes at a certain time period t (for example, 9:00 AM) are estimated using the people flow values ​​of four base meshes. That is, in the embodiment, people flow values ​​of some meshes (base meshes) from which data can be collected more easily are used to estimate people flow values ​​of the other meshes (surrounding meshes).

[0014] As described above, the technology of constructing an estimation model for each mesh may impose an excessive load on the model construction. In the embodiment, for example, the following functions are provided to more efficiently construct an estimation model for estimating the attribute of one of multiple regions. (F1) Using the people flow values ​​of the surrounding meshes, the surrounding meshes are classified (clustered) into multiple clusters. (F2) For each cluster, an estimation model is constructed that estimates the pedestrian flow values ​​of surrounding meshes from the pedestrian flow values ​​of the base mesh.

[0015] The above function can reduce the time required to build an estimation model (learning time) and the processing costs during learning, such as the amount of learning data used for learning. In other words, it is possible to build a model that estimates people flow values ​​more efficiently.

[0016] (First embodiment) 2 is a block diagram showing an example of the configuration of the information processing device 100 according to the first embodiment. As shown in FIG. 2, the information processing device 100 includes a storage unit 120, an acquisition unit 101, a generation unit 102, a construction unit 103, an estimation unit 104, and an output control unit 105.

[0017] The storage unit 120 stores various information used by the information processing device 100. For example, the storage unit 120 stores surrounding people flow data 121, base point people flow data 122, cluster information 123, and model information .

[0018] The surrounding people flow data 121 is data including people flow values ​​acquired for each surrounding mesh. Fig. 3 is a diagram showing an example of the data structure of the surrounding people flow data 121. As shown in Fig. 3, the surrounding people flow data 121 has a data structure in which the date and time are associated with the people flow value for each surrounding mesh.

[0019] In the following, the number of surrounding meshes is set to N (N is an integer equal to or greater than 2). The surrounding mesh IDs, which are identification information for identifying the N surrounding meshes, are represented by p1 to pN. The surrounding people flow data 121 is represented in a table format that stores the people flow values ​​for time period t in the N surrounding meshes for the learning period T. The representation format of the surrounding people flow data 121 is not limited to the table format, and may be any other format such as CSV format.

[0020] In the example of surrounding people flow data 121 in Figure 3, the first column, with the date and time as the column name, describes the dates and times included in the learning period T, and the second column and subsequent columns, with each surrounding mesh ID as the column name, describe the people flow values ​​of the surrounding meshes identified by the surrounding mesh IDs for the corresponding dates and times.

[0021] The date and time are written in the format "YYYY / MM / DD hh:mm:ss", for example. In this case, the date and time indicate the start time of the period (count period) during which people flow values ​​are measured (counted). The end time of the count period is set to immediately before the start time of the next line. Note that although this embodiment assumes equal time resolution, the date and time do not have to be equal intervals. If the intervals are not equal, the date and time are expressed as a start time and an end time, for example, "YYYY / MM / DD hh:mm:ss~YYYY / MM / DD hh:mm:ss".

[0022] Returning to the explanation of Figure 2, the base point people flow data 122 is data that includes people flow values ​​acquired at each base point mesh. Figure 4 is a diagram showing an example of the data structure of the base point people flow data 122. As shown in Figure 4, the base point people flow data 122 has a data structure that associates date and time with people flow values ​​for each base point mesh.

[0023] In the following, the number of base point meshes is M (M is an integer equal to or greater than 1). Furthermore, the base point mesh IDs, which are identification information for identifying the M base point meshes, are represented by c1 to cM. The base point people flow data 122 has the same data structure as the surrounding people flow data 121, except that the base point mesh IDs are used instead of the surrounding mesh IDs.

[0024] Returning to the explanation of Fig. 2, the cluster information 123 is information indicating the clusters generated by the generation unit 102. Details of the cluster information 123 will be described later. The model information 124 is information indicating the estimation model constructed by the construction unit 103. For example, the model information 124 is information in which a cluster ID, which is identification information for identifying a cluster, is associated with information on the estimation model constructed for the cluster.

[0025] The storage unit 120 can be configured from any commonly used storage medium, such as a flash memory, a memory card, a RAM (Random Access Memory), an HDD (Hard Disk Drive), or an optical disk.

[0026] Some or all of the data stored in the memory unit 120 (surrounding people flow data 121, base point people flow data 122, cluster information 123, model information 124) may be stored in physically different storage media, or in different storage areas of the same physically stored medium.

[0027] The acquisition unit 101 acquires various types of information used in the information processing device 100. For example, the acquisition unit 101 acquires data used when learning (constructing) an estimation model and data used when making an estimation using the constructed estimation model.

[0028] The data used during learning is, for example, the surrounding people flow data 121 and the base point people flow data 122 obtained in each time period of the learning period T. The data used during estimation is, for example, data indicating the people flow value at the base point mesh obtained in the time period to be estimated.

[0029] Below, we define the pedestrian flow value of the surrounding meshes identified by N surrounding mesh IDs (p1, p2, . . . , pN) in a certain time period t as {x p1 t ,x p2 t ,···,x pN t}. Also, the people flow value of the base mesh identified by M base mesh IDs (c1, c2, . . . , cM) in a certain time period t is expressed as {x c1 t ,x c2 t ,···,x cM t}.

[0030] The acquisition unit 101 may acquire information by any method, for example, a method of receiving information from an external device via a network, or a method of reading information from a storage medium.

[0031] The generation unit 102 generates a plurality of clusters for classifying a plurality of surrounding meshes using surrounding people flow data 121, which is the people flow values ​​of a plurality of surrounding meshes. The people flow values ​​of the plurality of surrounding meshes included in the surrounding people flow data 121 correspond to a plurality of attribute data (second attribute data) that represent the attributes of each of the surrounding meshes. In the following, it is assumed that the generation unit 102 generates K clusters (K is an integer of 2 or more).

[0032] For example, the generation unit 102 performs clustering so that multiple surrounding meshes with similar people flow values ​​are classified into the same cluster, and generates K clusters. The generation unit 102 stores cluster information 123, which is information about the generated clusters, in the storage unit 120.

[0033] The construction unit 103 constructs an estimation model for each of the generated clusters. The estimation model is a model that inputs base point people flow data and outputs attribute data representing the attributes of the cluster (hereinafter, referred to as cluster attribute data).

[0034] For example, for each of K clusters, the construction unit 103 constructs an estimation model for the cluster using training data including the base point people flow data 122 for training and the people flow values ​​(surrounding people flow data) of the surrounding meshes classified into the cluster. Hereinafter, the set F of K estimation models is defined as {f1, f2, . . . , f K}. f k represents an estimation model for the k-th cluster (k is an integer satisfying 1≦k≦K). The constructing unit 103 stores in the storage unit 120 model information 124, which is information on each constructed estimation model.

[0035] The estimation unit 104 executes an estimation process to estimate people flow values ​​of surrounding meshes using the constructed estimation model. For example, for a surrounding mesh designated as an estimation target among multiple surrounding meshes, the estimation unit 104 inputs base point people flow data into an estimation model constructed for the cluster into which the surrounding mesh is classified, and obtains cluster attribute data output by the estimation model. The estimation unit 104 estimates the obtained cluster attribute data as the people flow value of the designated surrounding mesh.

[0036] The surrounding meshes to be estimated are specified by, for example, the surrounding mesh ID. In the following, the surrounding mesh IDs to be estimated are designated as p q (q is an integer that satisfies 1≦q≦N). q The cluster into which the surrounding mesh is classified can be specified by the cluster information 123. In addition, the estimation model constructed for the specified cluster can be acquired from the model information 124.

[0037] The output control unit 105 controls the output of various information used in the information processing device 100. For example, the output control unit 105 outputs the result of the estimation process performed by the estimation unit 104. Any method may be used to output the information, and examples of applicable methods include displaying the information on a display device and transmitting the information to an external device via a network.

[0038] At least a part of each of the above units (acquisition unit 101, generation unit 102, construction unit 103, estimation unit 104, and output control unit 105) may be realized by one or more processing units. Each of the above units is realized, for example, by one or more processors. For example, each of the above units may be realized by having a processor such as a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit) execute a program, that is, by software. Each of the above units may be realized by a processor such as a dedicated IC (Integrated Circuit), that is, by hardware. Each of the above units may be realized by a combination of software and hardware. When multiple processors are used, each processor may realize one of the units, or may realize two or more of the units.

[0039] The information processing device 100 may be physically configured as one device, or may be physically configured as multiple devices. For example, the information processing device 100 may be constructed in a cloud environment. Furthermore, each unit in the information processing device 100 may be distributed across multiple devices. For example, the information processing device 100 (information processing system) may be configured to include a device (e.g., a learning device) having functions (e.g., a generation unit 102, a construction unit 103) necessary for constructing (learning) an estimation model, and a device (e.g., an estimation device) having functions (e.g., an estimation unit 104) necessary for estimation processing using the estimation model.

[0040] Next, a description will be given of a model learning process performed by the information processing apparatus 100 according to the first embodiment. Fig. 5 is a flowchart showing an example of the model learning process according to the first embodiment.

[0041] In the model learning process, first, the generation unit 102 executes a cluster generation process (step S101). In the cluster generation process, clusters related to surrounding meshes are generated using people flow value data of the surrounding meshes (surrounding people flow data 121). Next, the construction unit 103 executes a model construction process (step S102). In the model construction process, an estimated model of each cluster generated by the generation unit 102 is constructed using people flow value data of the base mesh (base point people flow data 122) and people flow value data of the surrounding meshes (surrounding people flow data 121).

[0042] Next, the cluster generation process in step S101 will be described in detail. Fig. 6 is a flowchart showing an example of the cluster generation process.

[0043] The generation unit 102 reads out surrounding people flow data 121 from, for example, the storage unit 120 (step S201). The generation unit 102 clusters surrounding meshes using the read out surrounding people flow data 121 (step S202). For example, the generation unit 102 clusters the surrounding meshes into K clusters based on the similarity of the time series data of people flow values ​​of the surrounding meshes. The generation unit 102 stores cluster information 123 indicating the correspondence between the generated clusters and the surrounding meshes in the storage unit 120.

[0044] FIG. 7 is a diagram showing an example of the data structure of the cluster information 123. The cluster information 123 is information indicating which of the K clusters N surrounding meshes belong to. In the example of FIG. 7, the cluster information 123 is expressed in a table format with a mesh ID and a cluster ID as column names. The mesh ID column is set to the surrounding mesh ID, which is the mesh ID of the surrounding mesh. The cluster ID column is set to the cluster ID (CL1, CL2, ..., CLK) of the cluster into which the surrounding mesh identified by the surrounding mesh ID in the same row is classified.

[0045] Next, the clustering process in step S202 will be described in detail below. Fig. 8 is a flowchart showing an example of the clustering process.

[0046] The generation unit 102 initializes a distance matrix and indexes i and j (step S301). The distance matrix is ​​an N-row, N-column matrix (N×N matrix) that represents the similarity between two surrounding meshes. The indexes i and j are integers between 1 and N. The generation unit 102 acquires a pair of unprocessed indexes i and j (j≠i) (step S302).

[0047] The generation unit 102 generates the i-th mesh (the surrounding mesh ID is p i ) and the jth (the surrounding mesh ID is p j The similarity (i, j) between the meshes surrounding the mesh (i, j) and the mesh (i, j) is calculated (step S303). The similarity (i, j) is expressed by, for example, the following formula (1). Similarity (i,j)= Sim({x pi t1 ,x pi t2 ,···,x pi tT}, {x pj t1 ,x pj t2 ,···,x pj tT}) ···(1)

[0048] t1, t2, . . . , tT represent the time points included in the specified learning period. pq t1 ~x pq tT The surrounding mesh ID is p q represents the people flow value of the surrounding mesh at times t1 to tT. Sim may be any function that calculates the similarity of time series data, but it is a function that calculates similarity using, for example, Euclidean distance, DTW (Dynamic Time Warping), autocorrelation, and Pearson correlation. Note that the similarity when i=j is set to 0.

[0049] The generation unit 102 stores the calculated similarity in the i-th row and j-th column of the distance matrix (step S304). The generation unit 102 determines whether all combinations of i and j have been processed (step S305). If all combinations of i and j have not been processed (step S305: No), the generation unit 102 returns to step S302 and repeats the process for the next unprocessed pair.

[0050] When all combinations of i and j have been processed (step S305: Yes), the generating unit 102 generates K clusters using the distance matrix (step S306).

[0051] For example, the generation unit 102 performs hierarchical clustering using a distance matrix and a preset number of clusters K to generate K clusters. Note that the clustering method is not limited to hierarchical clustering, and any other clustering method may be used. For example, a general clustering method for time-series data, such as the k-means method, may be used. The number of clusters K may also be set using the elbow method or the like. Furthermore, the generation unit 102 may normalize the people flow values ​​of each surrounding mesh before performing the clustering process. Normalization is, for example, a process of converting the average value to 0 and the standard deviation to 1.

[0052] Next, the details of the model construction process in step S102 will be described. Fig. 9 is a flowchart showing an example of the model construction process. In the model construction process, an estimated model showing the relationship between the base mesh and the surrounding meshes is constructed for each of the K clusters. As a result, K estimated models {f1, f2,..., f K} is obtained.

[0053] The construction unit 103 reads out the cluster information 123, the surrounding people flow data 121, and the base point people flow data 122 from the storage unit 120 (step S401). The construction unit 103 identifies an unprocessed cluster (step S402). The construction unit 103 generates learning data including the people flow values ​​of the surrounding meshes included in the identified cluster and the people flow value of the base point mesh (base point people flow data 122) (step S403). The people flow value of the surrounding meshes included in the identified cluster can be obtained, for example, from the surrounding people flow data 121 as the people flow value corresponding to the surrounding mesh ID of the surrounding mesh included in the identified cluster.

[0054] The construction unit 103 constructs an estimation model using the generated learning data (step S404). For example, the people flow value of the surrounding meshes belonging to the k-th cluster (cluster ID is CLk) is expressed as x k t Expressed as x k t and the pedestrian flow value of the base mesh {x c1 t ,x c2 t ,···,x cM t} is the relationship between the estimated model f k Using this, it is expressed by the following equation (2). x k t =f k (x c1 t ,x c2 t ,···,x cM t ) ···(2)

[0055] The construction unit 103 constructs the estimation model f kThe information is stored in the storage unit 120 as model information 124 (step S405). Similar to the clustering process, before executing the model construction process, the construction unit 103 may normalize the pedestrian flow values of each peripheral mesh. In this case, the construction unit 103 stores parameters related to normalization, for example, the average value and standard deviation of the original pedestrian flow values, in the storage unit 120 (model information 124). Thereby, it becomes possible to execute the same normalization even in the estimation process using the estimation model.

[0056] The construction unit 103 determines whether all clusters have been processed (step S406). If not all clusters have been processed (step S406: No), the construction unit 103 returns to step S202 and repeats the process for the next unprocessed cluster. If all clusters have been processed (step S406: Yes), the model construction process ends.

[0057] Next, a specific example of the model construction process will be described. Assume that L (where L is an integer satisfying 1 ≦ L < N) peripheral meshes belong to the k-th cluster. The pedestrian flow values of the L peripheral meshes are represented as {x pk1 t1 ,x pk1 t2 ,···,x pk1 tT ,···,x pkl t1 ,x pkl t2 ,···,x pkl tT ,···,x pkL t1 ,x pkL t2 ,···,x pkL tT )} (where l is an integer satisfying 1 ≦ l ≦ L). The construction unit 103 generates learning data including such pedestrian flow values of the peripheral meshes and the pedestrian flow values {x c1 t ,x c2 t ,···,x cM t} of the base point mesh.

[0058] Fig. 10 is a diagram showing an example of training data. Fig. 10 shows an example of training data 1001 in table format. In the training data 1001, the date and time are written in the first column, the surrounding mesh IDs are written in the second column, and the people flow values ​​of the surrounding meshes (surrounding people flow data) are written in the third column. The column names from the fourth column onwards are set to each of multiple base mesh IDs, and the people flow values ​​of the base meshes identified by the base mesh IDs are written.

[0059] The construction unit 103 associates the people flow values ​​of the surrounding meshes and the people flow values ​​of the base mesh with the date and time as a key to create learning data. For example, the people flow value of the surrounding meshes of a certain row LA is described as the people flow value obtained in the surrounding mesh identified by the surrounding mesh ID of row LA during the same time period as the date and time of row LA. In the fourth and subsequent columns of row LA, the people flow value obtained in the base mesh identified by the base mesh ID of each column during the same time period as the date and time of row LA is described.

[0060] The estimation model constructed in the model construction process may be of any type. In FIG. 10, the estimation model f k As shown in Figure 10, the multiple regression model calculates the pedestrian flow value {x c1 t ,x c2 t ,···,x cM t} and the weight of each base mesh {w c1,k t ,w c2,k t ,···,w cN,k t ,w 0,k t} and the linear combination 1012 of the surrounding meshes to calculate the pedestrian flow value 1011 (x k t ) is a model that represents w 0,k t is a constant weight.

[0061] Any method may be used to train the estimation model using the training data. For example, the construction unit 103 may use a training method using the least squares method or a training method using an optimization technique including regularization, such as an Elastic Net.

[0062] Next, a description will be given of the estimation process performed by the estimation unit 104. Fig. 11 is a flowchart showing an example of the estimation process according to the first embodiment.

[0063] The estimation unit 104 calculates p, which is the surrounding mesh ID of the surrounding mesh to be estimated. q The estimation unit 104 acquires the surrounding mesh ID p q The estimation unit 104 acquires the cluster ID of the cluster to which the cluster belongs from the cluster information 123 (step S502). The estimation unit 104 acquires information on the estimation model constructed for the cluster identified by the acquired cluster ID from the model information 124 (step S503).

[0064] The estimation unit 104 acquires the people flow value of the base mesh at date and time s as base mesh people flow data for estimation (step S504). Note that date and time s is, for example, a date and time when there are no people flow values ​​in the surrounding meshes. The estimation unit 104 inputs the acquired people flow value of the base mesh into the estimation model, and calculates the people flow value output by the estimation model based on the surrounding mesh ID p q The pedestrian flow value of the surrounding mesh is estimated as (step S505).

[0065] The output control unit 105 outputs the estimation result by the estimation unit 104 (step S506), and the estimation process ends.

[0066] In this way, in the first embodiment, an estimation model is constructed for each of multiple clusters into which surrounding meshes are classified using people flow values. This reduces the processing cost during learning, that is, it becomes possible to construct a model that estimates people flow values ​​more efficiently.

[0067] (Second embodiment) In the first embodiment, clustering is performed using the people flow values ​​of the surrounding meshes, whereas in the second embodiment, clustering is performed using information other than the people flow values.

[0068] 12 is a block diagram showing an example of the configuration of an information processing device 100-2 according to the second embodiment. As shown in FIG. 12, the information processing device 100-2 includes a storage unit 120-2, an acquisition unit 101, a generation unit 102-2, a construction unit 103, an estimation unit 104, and an output control unit 105.

[0069] In the second embodiment, the functions of the storage unit 120-2 and the generation unit 102-2 are different from those in the first embodiment. The other configurations and functions are the same as those in the block diagram of the information processing device 100 in the first embodiment shown in FIG. 2, so the same reference numerals are used and the description thereof will be omitted here.

[0070] The storage unit 120-2 differs from the storage unit 120 of the first embodiment in that it further stores surrounding mesh information 125-2.

[0071] Figure 13 is a diagram showing an example of the data structure of the surrounding mesh information 125-2. In the example of Figure 13, the surrounding mesh information 125-2 is represented in a table format with column names of mesh ID, land classification, latitude, longitude, green area ratio, and the mesh ID of the nearest base point. The surrounding mesh information 125-2 (land classification, latitude, longitude, green area ratio, and the mesh ID of the nearest base point) corresponds to area attribute data that represents attributes (features) other than the people flow value (amount of objects) for the surrounding meshes.

[0072] Land classification represents a classification based on land use and rights, and corresponds to information describing the use of surrounding meshes. Examples of land classification include residential areas, industrial areas, rice fields, and parks and green spaces. Land classification may also be used as a "use zone" defined in Article 8 of the City Planning Act, such as commercial districts, business districts, residential districts, industrial districts, and other areas. Latitude and longitude correspond to information describing the location of surrounding meshes. The latitude and longitude represent representative values ​​of the latitude and longitude of each mesh, such as the latitude and longitude of the mesh's center point. The green space area ratio represents the proportion of green space within a mesh. The green space area ratio can be calculated using land classification and satellite image data. The mesh ID of the closest base point represents the mesh ID of the base point closest to the surrounding mesh.

[0073] The generation unit 102-2 generates a plurality of clusters using the surrounding mesh information 125-2 together with the surrounding people flow data 121 (corresponding to the quantity data) as attribute data. For example, the generation unit 102-2 generates a plurality of clusters using the surrounding people flow data 121 for each value of the area attribute data specified among the area attribute data.

[0074] Next, the clustering process of this embodiment will be described in detail. Fig. 14 is a flowchart showing an example of the clustering process of the second embodiment. In Fig. 14, the k-means method is used as the clustering method.

[0075] The generating unit 102-2 reads out the surrounding mesh information 125-2 from the storage unit 120-2 (step S601). The generating unit 102-2 generates a parent cluster ID using a preset category item (step S602).

[0076] The preset category items correspond to designated area attribute data among the area attribute data. For example, area attribute data whose elements are not continuous values ​​but are expressed as categories, such as land classification and the nearest base point mesh ID, are set as category items. The generation unit 102-2 generates a parent cluster ID using the elements of such category items. A parent cluster is a cluster corresponding to the elements of a category item. The parent cluster ID is identification information that identifies a parent cluster.

[0077] For example, if land division is set as a category item, the generation unit 102-2 generates a parent cluster ID such as "residential area_cM1" for the "residential area" element of the land division. The generation unit 102-2 generates other parent cluster IDs in the same way for the other elements of the land division.

[0078] The generation unit 102-2 acquires unprocessed parent clusters from among the multiple parent clusters corresponding to the multiple generated parent cluster IDs (step S603). The generation unit 102-2 acquires unprocessed surrounding meshes belonging to the acquired parent cluster (step S604). For example, if the parent cluster ID is "residential area_cM1", the generation unit 102-2 acquires surrounding meshes that have the parent cluster "residential area" in their land classification from the surrounding mesh information 125-2. The generation unit 102-2 maps the acquired surrounding meshes into a 3-dimensional space (step S605).

[0079] The generation unit 102-2 determines whether all the surrounding meshes have been processed (step S606). If all the surrounding meshes have not been processed (step S606: No), the generation unit 102-2 returns to step S604 and repeats the process for the next surrounding mesh.

[0080] If all the surrounding meshes have been processed (step S606: Yes), the generation unit 102-2 performs clustering by the k-means method using the dimensional space (step S607). Note that the k-means method can also be applied to elements represented by categories by using binary variable processing.

[0081] The generation unit 102-2 generates a cluster ID from the parent cluster ID and the cluster generated in step S607 (step S608). For example, suppose that U clusters (U is an integer equal to or greater than 2) are generated by clustering the surrounding meshes whose parent cluster ID is "residential_cM1." In this case, the generation unit 102-2 generates U cluster IDs such as {residential_cM1_1, residential_cM1_2, . . . , residential_cM1_U}.

[0082] The generation unit 102-2 determines whether all parent clusters have been processed (step S609). If all parent clusters have not been processed (step S609: No), the generation unit 102-2 returns to step S603 and repeats the process for the next unprocessed parent cluster. If all parent clusters have been processed (step S609), the clustering process ends. Note that in this embodiment, the number of clusters generated is the sum of the number of clusters generated for each parent cluster for all parent clusters.

[0083] The generation unit 102-2 may perform clustering processing using information indicating the type of people flow data (type information). The type information is, for example, information indicating whether the people flow data is for a holiday or a weekday. In this case, the generation unit 102-2 performs clustering processing for each type of type information using the people flow data of the type indicated by the type information. For example, the generation unit 102-2 performs clustering processing twice, once for holidays and once for weekdays.

[0084] In this case, the construction unit 103 constructs an estimation model for each type of information and for each of a plurality of clusters, thereby making it possible to use different estimation models for holidays and weekdays.

[0085] As described above, the information processing device of the second embodiment can perform clustering using attribute data other than people flow values.

[0086] (Third embodiment) In the third embodiment, an example will be described in which a graph neural network (hereinafter referred to as GNN) is adopted as the estimation model.

[0087] 15 is a block diagram showing an example of the configuration of an information processing device 100-3 according to the third embodiment. As shown in FIG. 15, the information processing device 100-3 includes a storage unit 120-3, an acquisition unit 101, a generation unit 102, a construction unit 103-3, an estimation unit 104, and an output control unit 105.

[0088] In the second embodiment, the functions of the storage unit 120-3 and the construction unit 103-3 are different from those in the first embodiment. The other configurations and functions are the same as those in the block diagram of the information processing device 100 in the first embodiment shown in FIG. 2, so the same reference numerals are used and the description thereof will be omitted here.

[0089] The storage unit 120-3 differs from the storage unit 120 of the first embodiment in that it further stores base point mesh information 126-3. The base point mesh information 126-3 is information relating to M base point meshes.

[0090] Fig. 16 is a diagram showing an example of the data structure of the base point mesh information 126-3. Fig. 16 shows an example of the base point mesh information 126-3, which is an M x M adjacency matrix indicating the connection relationships of M base point meshes.

[0091] For example, when two base meshes with base mesh IDs c1 and cM are adjacent to each other, the elements of the corresponding row and column of the adjacency matrix will be 1. Conversely, when they are not adjacent to each other, the elements of the corresponding row and column of the adjacency matrix will be 0. The element values ​​are not limited to 0 or 1. For example, when they are adjacent to each other, a value such as the distance between the two base meshes may be set as the element value.

[0092] Two base meshes being adjacent to each other means, for example, that the two base meshes are adjacent (connected) in a network that shows the mutual connection relationships of multiple base meshes. The network showing the connection relationship may be, for example, a transportation network such as a railroad network or a road network. The adjacent relationship is not limited to a relationship determined using a network. For example, the two base meshes that are closest in distance may be defined as being adjacent to each other.

[0093] The construction unit 103-3 generates an estimation model that is a GNN. A GNN corresponds to a model that uses an adjacency matrix.

[0094] FIG. 17 is a diagram showing an overview of model construction using a GNN. The process up to generation of training data is the same as in the first embodiment. In a GNN, base mesh information 126-3, i.e., an adjacency matrix, is used when constructing a model. In the example of FIG. 17, an estimation model with a network structure having M nodes in the input layer and one node in the output layer is used. In a GNN, information on each node included in each layer, such as the intermediate layer and the output layer, is processed using the adjacency matrix (base mesh information 126-3).

[0095] In the estimation model in Fig. 17, to associate one base mesh with one node, a vector with M elements is input to the node that inputs the people flow value of the mth (m = 1, , M) base mesh. The mth element of the vector input to the mth node is set to the people flow value, and the other elements are set to 0.

[0096] Also, a vector with M elements is generated in the output layer. The construction unit 103-3 compares the output vector with a vector in which the people flow values ​​of the surrounding mesh to be learned are set for each of the M elements, calculates the error required for learning, and learns an estimation model to minimize the error. The m-th element x^ of the output vector is pim t represents the estimated value of the people flow value of the surrounding mesh whose surrounding mesh ID is pi.

[0097] 17 is used for estimation, a vector having M elements is output, and the estimation unit 104 uses a representative value of the M elements of the vector as an estimated value. The representative value may be, for example, the mean value, median, first quartile, or third quartile of the M elements.

[0098] In this way, in the third embodiment, the GNN can be used as an estimation model.

[0099] (Fourth embodiment) In the above embodiment, it is assumed that the plurality of meshes are classified into either base meshes or peripheral meshes. In the fourth embodiment, a function for classifying the plurality of meshes into either base meshes or peripheral meshes is provided.

[0100] Fig. 18 is a block diagram showing an example of the configuration of an information processing device 100-4 according to the fourth embodiment. As shown in Fig. 18, the information processing device 100-4 includes a storage unit 120, an acquisition unit 101, a generation unit 102, a construction unit 103, an estimation unit 104, an output control unit 105, and a classification unit 106-4.

[0101] The fourth embodiment differs from the first embodiment in that a classification unit 106-4 is added. Other configurations and functions are the same as those of the information processing device 100 of the first embodiment shown in FIG. 1, and therefore the same reference numerals are used and the description thereof will be omitted.

[0102] The classification unit 106-4 classifies the plurality of meshes into one or more base meshes and a plurality of peripheral meshes based on predetermined criteria, such as the following criteria: (R1) A criterion indicating whether predetermined facilities are installed. The predetermined facilities are facilities where the pedestrian flow value is expected to be large, such as a station. (R2) A criterion for indicating whether the amount of an object present in a mesh is greater than that of other meshes. The amount of an object being greater than that of other meshes means, for example, that the amount of the object is greater than a threshold, that the object is contained in a certain number of meshes in order of the amount of the object being large, or that the object is contained in a certain percentage of meshes in order of the amount of the object being large.

[0103] When the criterion (R1) is used, the classification unit 106-4 classifies, for example, a mesh in which a predetermined facility is installed as a base mesh, and classifies a mesh in which the facility is not installed as a peripheral mesh.

[0104] When the (R2) criterion is used, the classification unit 106-4 classifies, for example, a mesh whose number of objects (people flow value) is larger than that of other meshes as a base mesh, and classifies a mesh whose number of objects is not larger than that of other meshes as a surrounding mesh. In this case, the classification unit 106-4 performs the classification process using input data including the people flow value of each mesh. When people flow values ​​for multiple time periods are obtained, a representative value of the people flow value in each mesh may be compared with a threshold. The representative value may be, for example, the mean, median, first quartile, or third quartile of M elements.

[0105] Fig. 19 is a diagram showing an example of the data structure of input data used in the classification process. The input data is data containing people flow values ​​acquired in each of multiple meshes. As shown in Fig. 19, the input data has a data structure in which the date and time are associated with the people flow value for each of multiple meshes.

[0106] The input data is represented in a table format that stores people flow values ​​for time period t in each of multiple meshes for the learning period T. The representation format of the input data is not limited to the table format, and may be any other format such as CSV format.

[0107] The input data can also be interpreted as corresponding to the people flow data before being divided into the surrounding people flow data 121 and the base point people flow data 122. In other words, the input data may be configured so that data including people flow values ​​of meshes classified as base point meshes by the classification unit 106-4 is used as the base point people flow data 122, and data including people flow values ​​of meshes classified as peripheral meshes is used as the peripheral people flow data 121.

[0108] In Figure 19, for convenience of explanation, the mesh IDs are listed as c1, cM, and p1, which are the same as in the above embodiment, but this does not mean that the mesh is a base mesh (mesh ID is c1, cM) or a peripheral mesh (mesh ID is p1) at the time of input as input data. In other words, the mesh ID may be identification information that identifies all meshes before they are classified as base meshes or peripheral meshes.

[0109] Next, the classification process performed by classification unit 106-4 will be described with reference to a flowchart of FIG.

[0110] The classification unit 106-4 acquires people flow data corresponding to the input data (step S701). The classification unit 106-4 classifies the meshes included in the input data into either a base mesh or a surrounding mesh based on a preset criterion (step S702).

[0111] The classification unit 106-4 outputs the people flow data of the meshes classified as the base mesh as base point people flow data 122 (step S703). The classification unit 106-4 outputs the people flow data of the meshes classified as the surrounding mesh as surrounding people flow data 121 (step S704).

[0112] In this way, in the fourth embodiment, a plurality of meshes can be classified into either base meshes or peripheral meshes.

[0113] As described above, according to the first to fourth embodiments, an estimation model for estimating the attribute of one of a plurality of regions can be constructed more efficiently.

[0114] Next, the hardware configuration of the information processing apparatus according to the first to fourth embodiments will be described with reference to Fig. 21. Fig. 21 is an explanatory diagram showing an example of the hardware configuration of the information processing apparatus according to the first to fourth embodiments.

[0115] The information processing device of the first to fourth embodiments includes a control device such as a CPU (Central Processing Unit) 51, a storage device such as a ROM (Read Only Memory) 52 and a RAM (Random Access Memory) 53, a communication I / F 54 that connects to a network and communicates, and a bus 61 that connects each part.

[0116] The programs executed by the information processing apparatuses of the first to fourth embodiments are provided in advance in the ROM 52 or the like.

[0117] The programs executed by the information processing devices of the first to fourth embodiments may be configured to be provided as a computer program product by being recorded in an installable or executable file format on a computer-readable recording medium such as a CD-ROM (Compact Disk Read Only Memory), a flexible disk (FD), a CD-R (Compact Disk Recordable), or a DVD (Digital Versatile Disk).

[0118] Furthermore, the programs executed by the information processing apparatuses of the first to fourth embodiments may be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network. Also, the programs executed by the information processing apparatuses of the first to fourth embodiments may be provided or distributed via a network such as the Internet.

[0119] The programs executed by the information processing devices of the first to fourth embodiments can cause a computer to function as each unit of the information processing device described above. In this computer, the CPU 51 can read the programs from a computer-readable storage medium onto the main storage device and execute them.

[0120] A configuration example of the embodiment will be described below. (Configuration example 1) generating a plurality of clusters for classifying a plurality of second regions using a plurality of second attribute data representing attributes of each of a plurality of second regions other than one or more first regions included in the plurality of regions; constructing an estimation model, for each of the plurality of clusters, using training data including first attribute data for training and the second attribute data of the second regions classified into the clusters, which inputs one or more first attribute data indicating the attributes of each of the one or more first regions and outputs cluster attribute data indicating the attributes of the clusters; Processing section An information processing device comprising: (Configuration example 2) The processing unit inputting one or more pieces of first attribute data into the estimation model constructed for the cluster into which the second region is classified, for the second region designated as an estimation target among the plurality of second regions, and estimating the cluster attribute data output by the estimation model as attribute data indicating the attribute of the designated second region; The information processing device according to configuration example 1. (Configuration example 3) The processing unit generating a plurality of clusters such that a plurality of second regions having similar second attribute data are classified into the same cluster; The information processing device according to configuration example 1 or 2. (Configuration Example 4) the second attribute data includes quantity data representing the quantity of an object present in the corresponding second region, and region attribute data representing an attribute of the corresponding second region other than the quantity, the processing unit generates the plurality of clusters using the plurality of quantity data included in the plurality of second attribute data for each value of the specified region attribute data among the region attribute data. The information processing device according to any one of configuration examples 1 to 3. (Configuration Example 5) The area attribute data is data representing a part or all of the usage method, location, green area ratio, and nearest first area of ​​the second area. The information processing device according to configuration example 4. (Configuration Example 6) The processing unit classifying the plurality of regions into one or more first regions and a plurality of second regions based on predetermined criteria; 6. The information processing device according to any one of configuration examples 1 to 5. (Configuration Example 7) The criteria indicate whether or not predetermined equipment is installed; The processing unit Classifying an area where the equipment is installed into the first area, and classifying an area where the equipment is not installed into the second area. The information processing device according to configuration example 6. (Configuration Example 8) The criterion indicates whether the amount of objects present in the region is greater than that in other regions; The processing unit A region in which the amount is greater than other regions is classified into the first region, and a region in which the amount is not greater than other regions is classified into the second region. The information processing device according to configuration example 6. (Configuration Example 9) the estimation model is a graph neural network model using an adjacency matrix that represents a connection relationship between a plurality of nodes corresponding to each of the plurality of first regions; The information processing device according to any one of configuration examples 1 to 8. (Configuration Example 10) An information processing method executed by an information processing device, generating a plurality of clusters for classifying a plurality of second regions using a plurality of second attribute data representing attributes of each of a plurality of second regions other than one or more first regions included in the plurality of regions; constructing an estimation model for each of the plurality of clusters, the estimation model inputting one or more first attribute data indicating the attribute of each of the one or more first regions and outputting cluster attribute data indicating the attribute of the cluster, using training data including first attribute data for training and the second attribute data of the second regions classified into the cluster; An information processing method including: (Configuration Example 11) On the computer, generating a plurality of clusters for classifying a plurality of second regions using a plurality of second attribute data representing attributes of each of a plurality of second regions other than one or more first regions included in the plurality of regions; constructing an estimation model for each of the plurality of clusters, the estimation model inputting one or more first attribute data indicating the attribute of each of the one or more first regions and outputting cluster attribute data indicating the attribute of the cluster, using training data including first attribute data for training and the second attribute data of the second regions classified into the cluster; A program to execute.

[0121] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, and are also included in the scope of the invention and its equivalents as defined in the claims. [Explanation of symbols]

[0122] 100, 100-2, 100-3, 100-4 Information processing device 101 Acquisition Department 102, 102-2 generation section 103, 103-3 Construction Department 104 Estimation part 105 Output control section 106-4 Classification Department 120, 120-2, 120-3 Storage section 121 Surrounding pedestrian flow data 122 Base point flow data 123 Cluster Information 124 Model Information 125-2 Surrounding mesh information 126-3 Base point mesh information

Claims

1. generating a plurality of clusters for classifying the plurality of second regions using a plurality of second attribute data representing attributes of each of a plurality of second regions other than the one or more first regions included in the plurality of regions; constructing an estimation model, for each of the plurality of clusters, using training data including first attribute data for training and the second attribute data of the second regions classified into the clusters, the estimation model inputting one or more first attribute data indicating the attributes of each of the one or more first regions and outputting cluster attribute data indicating the attributes of the clusters; Processing section An information processing device comprising:

2. The processing unit one or more pieces of first attribute data are input to the estimation model constructed for the cluster into which the second region is classified, for the second region designated as an estimation target among the plurality of second regions, and the cluster attribute data output by the estimation model are estimated as attribute data indicating the attribute of the designated second region. The information processing device according to claim 1 .

3. The processing unit generating a plurality of clusters such that a plurality of second regions having similar second attribute data are classified into the same cluster; The information processing device according to claim 1 .

4. the second attribute data includes quantity data representing an amount of an object present in the corresponding second region, and region attribute data representing an attribute of the corresponding second region other than the quantity, the processing unit generates the plurality of clusters using the plurality of quantity data included in the plurality of second attribute data for each value of the specified region attribute data among the region attribute data. The information processing device according to claim 1 .

5. The area attribute data is data representing a part or all of the usage method, location, green area ratio, and nearest first area of ​​the second area. The information processing device according to claim 4 .

6. The processing unit classifying the plurality of regions into one or more first regions and a plurality of second regions based on predetermined criteria; The information processing device according to claim 1 .

7. The criteria indicate whether or not predetermined equipment is installed; The processing unit Classifying an area where the equipment is installed into the first area, and classifying an area where the equipment is not installed into the second area. The information processing device according to claim 6 .

8. The criterion indicates whether the amount of objects present in the region is greater than that in other regions; The processing unit A region in which the amount is greater than other regions is classified into the first region, and a region in which the amount is not greater than other regions is classified into the second region. The information processing device according to claim 6 .

9. the estimation model is a graph neural network model using an adjacency matrix that represents a connection relationship between a plurality of nodes corresponding to each of the plurality of first regions; The information processing device according to claim 1 .

10. An information processing method executed by an information processing device, generating a plurality of clusters for classifying a plurality of second regions using a plurality of second attribute data representing attributes of each of a plurality of second regions other than one or more first regions included in the plurality of regions; constructing, for each of the plurality of clusters, an estimation model that inputs one or more first attribute data indicating the attributes of each of the one or more first regions and outputs cluster attribute data that indicates the attributes of the cluster, using training data including first attribute data for training and the second attribute data of the second regions classified into the cluster; An information processing method including:

11. On the computer, generating a plurality of clusters for classifying a plurality of second regions using a plurality of second attribute data representing attributes of each of a plurality of second regions other than one or more first regions included in the plurality of regions; constructing, for each of the plurality of clusters, an estimation model that inputs one or more first attribute data indicating the attributes of each of the one or more first regions and outputs cluster attribute data that indicates the attributes of the cluster, using training data including first attribute data for training and the second attribute data of the second regions classified into the cluster; A program to execute.

Citation Information

Patent Citations

  • Dicyclohexylcarbonylanthracene and luminescent composition

    JP1980076838A