Information processing device, information processing method, and information processing program
The information processing apparatus addresses the issue of inaccurate pedestrian flow data by converting and controlling specific components using principal component analysis, ensuring data accuracy and user-specific relevance.
Patent Information
- Application Number
- JP2024006273
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-18
- Publication Date
- 2025-07-31
- Estimated Expiration
- 2044-01-18
AI Technical Summary
Conventional pedestrian flow data analysis methods fail to provide highly accurate data tailored to the user's purpose, often removing important features unintentionally during correction processes.
An information processing apparatus and method that converts time-series pedestrian flow data into function data, extracts components via principal component analysis, and controls specific components based on user needs, while retaining others, to generate customized pedestrian flow data.
Enables the generation of pedestrian flow data that accurately reflects user-specific requirements by selectively modifying and retaining relevant features, improving data accuracy and scenario estimation.
Smart Images

Figure 2025112148000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, and an information processing program.
Background Art
[0002] Conventionally, analysis methods for pedestrian flow data have been proposed.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, it can be said that there is room for improvement in the above conventional technology in terms of providing pedestrian flow data according to the user's purpose.
[0005] For example, in the above conventional technology, the position information of the user terminal is acquired, and the movement path of the user terminal is generated by arranging it in time series according to the time information indicating the timing when the position information is acquired. Based on the trip data obtained from the processed movement path, the number of samples (the number of user terminals for which position information is acquired) is expanded to the population of the estimation area to estimate the flow of people in the population of the target area.
[0006] From this, it can be said that the above conventional technology efficiently analyzes pedestrian flow data over a wide range. On the other hand, in the above conventional technology, for example, it is not possible to generate highly accurate pedestrian flow data in which some features included in the pedestrian flow data are controlled according to the user's purpose while other features are reflected as they are. That is, the above conventional technology does not always provide pedestrian flow data according to the user's purpose.
[0007] Therefore, the present invention proposes an information processing apparatus, an information processing method, and an information processing program capable of providing people flow data according to the purpose of a user.
Means for Solving the Problems
[0008] In order to solve the above problems, an information processing apparatus according to one aspect of the present invention includes a conversion unit that converts time-series people flow data, which is people flow data of the population in a predetermined range, into function data; an extraction unit that extracts a plurality of people flow components by performing principal component analysis on the function data; a people flow control unit that controls so that a predetermined people flow component among the plurality of people flow components is changed according to the purpose of the user; and a generation unit that generates people flow data to be provided to the user based on the plurality of people flow components including the changed people flow component.
Effects of the Invention
[0009] According to the present invention, it is possible to provide people flow data according to the purpose of the user.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Embodiments for Carrying Out the Invention
[0011] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In the present specification and drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant description is omitted.
[0012] One or more of the embodiments (including examples, modifications, and application examples) described below can be implemented independently. On the other hand, at least some of the plurality of embodiments described below may be implemented in appropriate combination with at least some of other embodiments. These multiple embodiments may include different novel features. Therefore, these multiple embodiments can contribute to solving different objectives or problems and can exhibit different effects.
[0013] (Embodiment) [1. Introduction] Currently, time-series data is utilized in various services. For example, time-series data from sensors may be used for various anomaly detections. For example, prediction services such as detection of disasters and accidents or detection of anomalies in railway operation status are realized based on time-series sensor data indicating human movement.
[0014] In addition, there are also services that predict how sales change according to temperature and the flow of people based on time-series sales data.
[0015] On the one hand, the pedestrian flow data is not only required to be utilized for prediction, but may also need to be corrected itself. Generally, it is considered that the pedestrian flow tends to be the same between the past and the present. Therefore, it is required to appropriately correct the pedestrian flow data according to the purpose.
[0016] For example, among the features included in this year's pedestrian flow data, there is a need to control specific features that are considered not to occur normally from past cases. To give an example, among the features included in the pedestrian flow data, for features corresponding to special population increases and decreases that are estimated not to occur this year (for example, a sharp decrease in population due to movement restrictions under a state of emergency, or a rapid increase in population due to the lifting of movement restrictions), there is a need to delete them from this year's pedestrian flow data.
[0017] Also, for example, when assuming that the pedestrian flow increases, or when assuming that the pedestrian flow decreases, there may be a case where not only predicting how the future pedestrian flow will be, but estimating a scenario according to the prediction result is required.
[0018] Therefore, in the present invention, a method for generating highly accurate pedestrian flow data by correcting the features of the pedestrian flow data so as to satisfy the needs required by the user for the pedestrian flow data is proposed. For example, as a result of changing the features included in the pedestrian flow data according to the user's purpose, there is a possibility that other important features may also be lost, and the obtained pedestrian flow data cannot be said to have good accuracy and is difficult for scenario estimation. Therefore, the purpose of the present invention is to generate highly accurate pedestrian flow data while changing some of the features included in the pedestrian flow data according to the user's purpose and reflecting the other features as they are without control.
[0019] [2. Pedestrian Flow Data] First, the flow-of-people data used in the information processing according to the proposed technology of the present invention (information processing according to the embodiment) will be described. The flow-of-people data referred to here indicates a transition matrix of the number of people who have moved from one range to another range between a certain specific timing (for example, date and time) and the next timing. For example, by applying principal component analysis (PCA: Principal Component Analysis) to the preprocessed multi-dimensional flow-of-people data, converting it into two dimensions, and plotting it, the features corresponding to the principal components can be visualized.
[0020] An example of the flow-of-people data will be described with reference to FIGS. 1 and 2. The flow-of-people data shown in FIGS. 1 and 2 is multi-dimensional (multivariable) data including a plurality of variables regarding the observation of the population in a predetermined period for each predetermined range, and has the concept of time series. That is, the flow-of-people data shown in FIGS. 1 and 2 is time-series flow-of-people data including a plurality of variables.
[0021] First, FIG. 1 will be described. FIG. 1 is a diagram (1) showing an example of the flow-of-people data according to the embodiment. The flow-of-people data shown in FIG. 1 is multi-dimensional data with relatively few variables, i.e., small-variable data.
[0022] An observation example for obtaining the flow-of-people data is shown in FIG. 1(a). As shown in FIG. 1(a), the flow-of-people data is observed in units of trip classes (Trip Class) in which regions are classified according to the moving distance, which is the distance that the user moves (trips). The trip class is an example of a predetermined range. In FIG. 1(a), as trip classes, a set of "Trip Class 1", "Trip Class 2", "Trip Class 3", and "Trip Class 4" (denoted as "Trip Class 1" ··· "Trip Class 4") is shown.
[0023] Also, as shown in Fig. 1(a), the pedestrian flow data is time-series data obtained by statistically processing the number of people (population) sequentially observed (for example, observed every second) within a given time period for each such time period. In Fig. 1(a), as an example of a given time period, the time periods of "03:00 - 10:59", "11:00 - 14:59", "15:00 - 18:59", and "19:00 - 26:59" are shown. Further, according to the example of Fig. 1(a), more specifically, the pedestrian flow data is a compilation of the time-series data of the population obtained for each time period on a daily basis over 274 days from July 1, 2020 to March 31, 2021.
[0024] Here, according to Fig. 1(a), the pedestrian flow data contains a plurality of variables. This point will be explained with reference to Fig. 1(b). For example, when the time-series data of the population obtained for each of the 4 types of trip classes for each of the 4 time periods is compiled on a daily basis over 274 days as the pedestrian flow data, such pedestrian flow data contains 4×4 = 16 variables. Specifically, there are 16 combinations of variables between 4 class variables ("Trip Class 1", "Trip Class 2", "Trip Class 3", and "Trip Class 4") and 4 time period variables (the time period of "03:00 - 10:59", the time period of "11:00 - 14:59", the time period of "15:00 - 18:59", and the time period of "19:00 - 26:59").
[0025] That is, as shown in Fig. 1(b), for each of the 16 variables, there is time-series pedestrian flow data for 274 days. In Fig. 1(b), "Pedestrian Flow Data DA11" is shown as the time-series pedestrian flow data corresponding to the class variable "Trip Class 1" and the time period variable "11:00 - 14:59". Explanation of the pedestrian flow data corresponding to each of the other 15 variable sets is omitted.
[0026] In this way, when the time-series data of the population obtained for each of the four time periods for each of the four trip classes was summarized daily over 274 days to obtain the pedestrian flow data, since such pedestrian flow data includes a relatively small number of 16 variables, it can be said to be multi-dimensional data with a small number of variables.
[0027] Note that how the pedestrian flow data is summarized is arbitrary, and it can also be obtained as multi-dimensional data including a large number of variables. An example showing this is Fig. 2. Fig. 2 is a diagram (2) showing an example of the pedestrian flow data according to the embodiment.
[0028] Fig. 2(a) shows an example of observation for obtaining the pedestrian flow data. As shown in Fig. 2(a), the pedestrian flow data is observed in units of meshes where the regions on the map are divided into a grid based on latitude and longitude. A mesh is an example of a predetermined range. Fig. 2(a) shows 2385 meshes from "mesh0001" to "mesh2385" (denoted as "mesh0001", "mesh0002", "mesh0003", "mesh0004", ··· "mesh2382", "mesh2383", "mesh2384", "mesh2385") as the meshes.
[0029] Also, similar to the example in Fig. 1(a), the pedestrian flow data is time-series data obtained by statistically processing the number of people (population) sequentially observed (for example, observed every second) within the corresponding time period for each predetermined time period.
[0030] Here, according to Fig. 2(a), the pedestrian flow data includes a plurality of variables. This point will be described with reference to Fig. 2(b). For example, when the time-series data of the population obtained for each of the 2385 types of meshes was summarized daily over 274 days to obtain the pedestrian flow data, such pedestrian flow data includes 2385 variables.
[0031] That is, as shown in Fig. 2(b), for each of the 2,385 mesh variables, there will be time-series pedestrian flow data for 274 days. In Fig. 2(b), "Pedestrian flow data X1" is shown as the time-series pedestrian flow data corresponding to the mesh variable "mesh0001". Explanation of the pedestrian flow data corresponding to each of the other 2,384 mesh variables is omitted.
[0032] In this way, when the time-series data of the population obtained for each of the 2,385 types of meshes is summarized daily over 274 days and used as pedestrian flow data, such pedestrian flow data can be said to be multi-variable multi-dimensional data because it contains a large number of variables, 2,385.
[0033] [3. Prior Art] Prior to explaining the proposed technology of the present invention, the comparative prior art will be explained. Fig. 3 is an explanatory diagram for explaining the prior art. In Fig. 3(a), a conceptual diagram of the prior art for correcting pedestrian flow data is shown. According to the example in Fig. 3(a), in the prior art, pedestrian flow data for each variable is acquired, and the features included in the pedestrian flow data are corrected according to the user's purpose by complementary processing (linear complementation, curve complementation).
[0034] Here, in Fig. 3(b), the prior art will be specifically explained using as an example the scenario of correcting the 274-day pedestrian flow data DA11 described in Fig. 1. For example, in the period "from July 1, 2020 to March 31, 2021" included in the pedestrian flow data DA11, the population is rapidly decreasing during the "T1 period". Also, in the period "from July 1, 2020 to March 31, 2021" included in the pedestrian flow data DA11, the population is also significantly decreasing during the "T2 period".
[0035] If the user can infer that such a population decline will not occur this year, the user may desire the pedestrian flow data DA11 in a state where the characteristics of population decline during the "T1 period" and the characteristics of population decline during the "T2 period" are removed. In such a case, in the prior art, through complementary processing, the characteristics of the "T1 period" and the characteristics of the "T2 period" are removed, and the pedestrian flow data DA111 is generated, which is corrected so that the data of the removed part becomes smooth.
[0036] According to such prior art, among the characteristics included in the pedestrian flow data DA11 during the "T1 period", not only the characteristic of population decline but also other important characteristics may be removed. Similarly, for the pedestrian flow data DA11 during the "T2 period", not only the characteristic of population decline but also other important characteristics may be removed. Therefore, it is hard to say that the pedestrian flow data DA111 reflects the user's needs, and there is room for improvement in providing pedestrian flow data according to the user's purpose.
[0037] 〔4. Proposed Technology〕 The idea proposed to improve the above problems of the prior art is the proposed technology of the present invention. FIG. 4 is an explanatory diagram for explaining the proposed technology. In FIG. 4(a), a conceptual diagram of the proposed technology for correcting pedestrian flow data is shown. According to the example in FIG. 4(a), in the proposed technology, the pedestrian flow data for each variable is acquired, and each of the acquired discrete pedestrian flow data is converted into function data that can be represented by a curve.
[0038] Also, by performing principal component analysis on each function-data pedestrian flow data, a plurality of pedestrian flow components are extracted. Further, in the proposed technology, among the plurality of pedestrian flow components, only the pedestrian flow components corresponding to the user's purpose are controlled according to this purpose.
[0039] For example, as a result of being able to infer from the contribution rate and the principal component loading amount when extracting a plurality of people flow components that the first principal component represents a trend, a user may have an objective of deleting the features of a predetermined period included in the first principal component. In such a case, in the proposed technique, a complementation process (linear complementation, curve complementation) based on the principal component score when extracting the first principal component is performed. Then, in the proposed technique, people flow data is reconstructed (restored) from the plurality of people flow components including the complemented principal component.
[0040] Here, also in FIG. 4(b), the proposed technique will be specifically described using as an example the scenario of correcting the people flow data DA11 for 274 days described with reference to FIG. 1. As described with reference to FIG. 3, among the period "from July 1, 2020 to March 31, 2021" included in the people flow data DA11, the population is rapidly decreasing during the "T1 period". Also, among the period "from July 1, 2020 to March 31, 2021" included in the people flow data DA11, the population is also significantly decreasing during the "T2 period".
[0041] Therefore, the user considers deleting these features from the first principal component that expresses the feature of the population decrease during the "T1 period" and the trend feature of the population decrease during the "T2 period". In such a case, in the proposed technique, only the feature of the population decrease among the features included in the people flow data DA11 during the "T1 period" is deleted from the first principal component, and also only the feature of the population decrease among the features included in the people flow data DA11 during the "T2 period" is deleted from the first principal component.
[0042] As a result, since the first principal component in a state where the features corresponding to the user's objective are deleted is obtained, in the proposed technique, the people flow data is reconstructed (restored) based on the first principal component in which control for deleting the feature of the population decrease has been performed and the other principal components that remain uncontrolled. FIG. 4(b) shows an example in which the people flow data DA112 is generated by the reconstruction.
[0043] According to the pedestrian flow data DA112, while the characteristics of population decrease that the pedestrian flow data DA11 has during the "T1 period" are removed, the characteristics corresponding to the principal components other than the first principal component are retained. Also, according to the pedestrian flow data DA112, the characteristics of population decrease that the pedestrian flow data DA11 has during the "T2 period" are likewise removed, but the characteristics corresponding to the principal components other than the first principal component are retained.
[0044] In this way, in the proposed technology, only the characteristics according to the user's purpose are controlled, while other characteristics are retained without being controlled. Therefore, compared with the prior art, it is possible to generate pedestrian flow data corrected with high accuracy.
[0045] [5. System Configuration] FIG. 5 is a diagram showing an example of the system according to the embodiment. In FIG. 5, as an example of the system according to the embodiment, a system 1 is shown. The information processing according to the embodiment (that is, the proposed technology of the present invention) is realized in the system 1.
[0046] As shown in FIG. 1, the system 1 includes a user device 10 and an information processing device 100. Also, the user device 10 and the information processing device 100 are communicably connected by wire or wirelessly via a network N. Further, the information processing device 100 operates according to the program according to the embodiment.
[0047] The user device 10 is an information processing terminal used by the user and corresponds to an edge computer. For example, the user device 10 is a smartphone, a wearable device, a tablet terminal, a notebook PC (Personal Computer), a desktop PC, a mobile phone, a PDA (Personal Digital Assistant), or the like.
[0048] An environment (application) that can access the information processing apparatus 100 may be introduced into the user device 10. The application referred to here may be a general-purpose application such as a browser, or may be a dedicated application for accessing the information processing apparatus 100.
[0049] Also, the user referred to here is a person who wants to edit the features included in the human flow data according to the purpose, and may be an administrator of the information processing apparatus 100, a service user who requests the provision of human flow data, or the like.
[0050] The information processing apparatus 100 is a server apparatus that performs information processing according to the proposed technology of the present invention described with reference to FIG. 4, and corresponds to a cloud computer. A detailed configuration example of the information processing apparatus 100 will be described later.
[0051] 〔6. Configuration of Information Processing Apparatus〕 The information processing apparatus 100 according to the embodiment will be described with reference to FIG. 6. FIG. 6 is a diagram showing a configuration example of the information processing apparatus 100 according to the embodiment. According to the example of FIG. 6, the information processing apparatus 100 includes an external interface unit 111, a management interface unit 112, and an input / output interface unit 113.
[0052] The information processing apparatus 100 also includes an input storage unit 120, a storage unit 130, and an output storage unit 140. According to the example of FIG. 6, the input storage unit 120 includes a human flow data storage unit 120a and a set value storage unit 120b. The storage unit 130 includes a grouped human flow storage unit 130a and a parameter storage unit 130b. The output storage unit 140 has a principal component human flow storage unit 140a.
[0053] Further, the information processing apparatus 100 includes a preprocessing unit 151, an operation unit 152, a clustering unit 153, a learning unit 154, an inference unit 155, an editing unit 156, and a generation unit 157. According to the example of FIG. 6, the learning unit 154 includes a function data conversion unit 154a and a PCA learning unit 154b. The inference unit 155 includes a principal component score calculation unit 155a and a principal component population flow conversion unit 155b. The editing unit 156 includes a score change unit 156a and a load amount change unit 156b. The generation unit 157 includes a principal component population flow correction unit 157a and a reconstruction unit 157b.
[0054] (External interface unit 111) The external interface unit 111 acquires time-series population flow data. The external interface unit 111 acquires time-series population flow data from an external device (for example, the user device 10). The time-series population flow data is raw data of the population flow that is not summarized using the variables as described in FIGS. 1 and 2.
[0055] (Management interface unit 112) When the administrator performs an input operation using the user device 10, the management interface unit 112 receives input information corresponding to the input operation. For example, the management interface unit 112 receives information for editing the characteristics of the population flow data from the administrator.
[0056] (Input / output interface unit 113) When the service user performs an input operation using the user device 10, the input / output interface unit 113 receives input information corresponding to the input operation. For example, the input / output interface unit 113 receives information for editing the characteristics of the population flow data from the service user.
[0057] Further, the input / output interface unit 113 outputs to the user device 10 the population flow data reconstructed based on the principal components edited according to the user's input information.
[0058] (Population flow data storage unit 120a) The crowd flow data storage unit 120a stores the crowd flow data obtained by preprocessing the time-series number of people data. For example, the crowd flow data storage unit 120a stores the crowd flow data reorganized by dividing the time-series number of people data by statistical methods (e.g., kshape) or qualitative criteria (e.g., a predetermined range such as a mesh).
[0059] (Setting value storage unit 120b) The setting value storage unit 120b stores the setting values for function data analysis, that is, for converting the crowd flow data into function data. The setting values may be input from an administrator via the management interface unit 112, for example. The setting values may be, for example, initial values used in converting the crowd flow data into function data in function data conversion.
[0060] (Grouped crowd flow memory unit 130a) The grouped crowd flow memory unit 130a stores the grouped crowd flow data, which is data obtained by clustering the crowd flow data. Here, using the example of Fig. 1(b), there will be 16 variables × 274 (total 4384) crowd flow data. Depending on the way of grouping according to the variables, there may be even more crowd flow data. Thus, when the number of crowd flow data is huge, it will take a lot of time for principal component analysis, and there is a risk that the features will be scattered and appropriate principal component analysis cannot be performed. For this reason, in this embodiment, the crowd flow data is clustered in a predetermined unit. For example, in this embodiment, the crowd flow data is clustered among meshes containing crowd flow data having the same (similar) features. Note that the clustering process does not necessarily have to be performed.
[0061] (Parameter memory unit 130b) The parameter storage unit 130b stores parameters used when learning a series of processes of an algorithm using principal component analysis. The principal component analysis algorithm is a machine learning algorithm for solving a maximization problem that maximizes the variance of data. More specifically, the algorithm for performing principal component analysis (functional principal component analysis) on function data is a machine learning algorithm for solving a maximization problem that maximizes the variance of composite variables (principal component scores) represented by the inner product of function data and weight functions.
[0062] (Principal component inflow storage unit 140a) The principal component inflow storage unit 140a stores inflow data (principal component inflow data) for each principal component (inflow component) extracted by performing principal component analysis on the inflow data. For example, the principal component inflow storage unit 140a stores data obtained by decomposing the inflow data for each principal component. Here, the extraction of the principal component means extracting new variables (principal components) from the original variables in principal component analysis. The principal component represents an axis that has been transformed in the direction of higher importance of information while retaining the information of the original variables. That is, the principal component (inflow component) summarizes the inflow data and is represented in a coordinate system. Therefore, the principal component inflow data referred to here is obtained by restoring the original function data (unit: person (population)) for each of a plurality of principal component functions obtained by summarization through principal component analysis.
[0063] (Preprocessing unit 151) The preprocessing unit 151 performs preprocessing of reorganizing by dividing the time-series population data using statistical methods (e.g., kshape) or qualitative criteria (e.g., a predetermined range such as a mesh). Also, the preprocessing unit 151 registers the inflow data obtained by the preprocessing in the inflow data storage unit 120a.
[0064] (Operation unit 152) When the operation unit 152 receives input information for clustering the pedestrian flow data through the management interface, it operates according to this input information. For example, the operation unit 152 acquires the pedestrian flow data from the pedestrian flow data storage unit 120a and operates the clustering unit 153 to cluster the acquired pedestrian flow data according to the input information.
[0065] (Clustering unit 153) The clustering unit 153 executes the above-described clustering process. For example, the clustering unit 153 clusters the pedestrian flow data among the meshes that include pedestrian flow data having similar (alike) characteristics based on the input information transmitted from the operation unit 152. In addition, the clustering unit 153 registers the pedestrian flow data for each class (grouped pedestrian flow data) in the grouped pedestrian flow storage unit 130a.
[0066] (Function data conversion unit 154a) The function data conversion unit 154a executes a function data conversion process for converting discrete pedestrian flow data into function data. For example, the function data conversion unit 154a acquires the set value from the set value storage unit 120b, acquires the grouped pedestrian flow data from the grouped pedestrian flow storage unit 130a, and converts each grouped pedestrian flow data for each class into function data based on the set value.
[0067] Here, FIG. 7 is a diagram in which the pedestrian flow data shown in FIG. 1 is arranged and plotted for each variable. In FIG. 7, graphs of 16 pieces of pedestrian flow data are shown according to 16 variables.
[0068] For example, assume that 16 variables × 274 pieces (a total of 4384 pieces) of pedestrian flow data as shown in FIG. 1(b) are generated by preprocessing. In such a case, the clustering unit 153 may classify the pedestrian flow data in mesh groups corresponding to 16 variables by a clustering process based on the relationship between the characteristics of the pedestrian flow data and the meshes. The function data conversion unit 154a converts each pedestrian flow data into function data based on the plot shown in FIG. 7.
[0069] (PCA Learning Unit 154b) Returning to FIG. 6, the PCA learning unit 154b extracts a plurality of human flow components by performing principal component analysis on the function data. For example, the PCA learning unit 154b calculates eigenvectors and eigenvalues and registers them in the parameter storage unit 130b. From this, the PCA learning unit 154b is a processing unit corresponding to the extraction unit. Also, the PCA learning unit 154b executes learning processing to maximize the variance of the data in the function data. Further, the PCA learning unit 154b registers in the parameter storage unit 130b the parameters that can maximize the variance. Note that in the machine learning algorithm by the PCA learning unit 154b, when learning the axis that can most represent the data from the learning data, weights for optimizing the loss function are searched for.
[0070] (Inference Unit 155) The inference unit 155 returns the data that has been subjected to principal component scoring to the original data using the eigenvectors. For example, the inference unit 155 acquires from the parameter storage unit 130b the data that has been subjected to principal component scoring and the eigenvectors. For example, when the eigenvectors are changed by the editing unit 156, the inference unit 155 calculates the principal component scores with the acquired eigenvectors replaced by the changed eigenvectors. Note that when the principal component scores are changed by the editing unit 156, the inference unit 155 replaces the calculated principal component scores with the changed principal component scores. Then, the inference unit 155 restores the original data based on the current principal component scores and the eigenvectors. Also, the inference unit 155 generates human flow data for each principal component using the plurality of human flow components extracted by performing principal component analysis on the function data. From this, the inference unit 155 is a processing unit corresponding to the extraction unit. Note that the inference unit 155 may calculate the principal component loadings based on the eigenvalues and eigenvectors calculated by the principal component score calculation unit 155a.
[0071] (Principal Component Score Calculation Unit 155a) The principal component score calculation unit 155a calculates the principal component scores based on the pedestrian flow data. For example, the principal component score calculation unit 155a calculates the principal component scores for each piece of function data (in the example of FIG. 7, 16 pieces of function data) in which the pedestrian flow data belonging to each mesh group is functionalized. The principal component scores are used for the calculation of the principal component pedestrian flow data. For example, the principal component score calculation unit 155a generates a covariance matrix by calculating the covariance between the variables included in the function data. Then, the principal component score calculation unit 155a calculates the eigenvalues and eigenvectors of the covariance matrix. For example, the principal component score calculation unit 155a can obtain parameters (for example, weights and principal component vectors) from the parameter storage unit 130b and calculate the principal component scores based on the obtained parameters. The calculation of the principal component scores is to calculate the position (principal component scores) on the axis, whereby the principal component pedestrian flow data corresponding to the purpose can be generated.
[0072] (Principal component pedestrian flow conversion unit 155b) The principal component pedestrian flow conversion unit 155b creates new data (principal component scores) by transforming the data using the eigenvectors. For example, the principal component pedestrian flow conversion unit 155b creates new data (principal component scores) by projecting the original data onto the eigenvectors. Thereby, the principal component pedestrian flow conversion unit 155b can transform the characteristics of the original pedestrian flow data having a plurality of variables into a two-dimensional graph.
[0073] Also, the principal component pedestrian flow conversion unit 155b registers the pedestrian flow data for each generated principal component in the principal component pedestrian flow storage unit 140a. For example, the original pedestrian flow data can be obtained by adding up all the principal component pedestrian flow data.
[0074] Here, FIG. 8 shows an example of the principal component pedestrian flow data. In FIG. 8, four principal components from the first principal component to the fourth principal component are extracted in a state where the 274-day pedestrian flow data DA11 described in FIG. 1 is converted into function data, and a scene where each principal component is converted into a two-dimensional graph is shown. Specifically, FIG. 8(a) shows a two-dimensional graph corresponding to the first principal component, and FIG. 8(b) shows a two-dimensional graph corresponding to the second principal component. Also, FIG. 8(c) shows a two-dimensional graph corresponding to the third principal component, and FIG. 8(d) shows a two-dimensional graph corresponding to the fourth principal component.
[0075] And based on the principal component loadings, for example, the user can infer which variables are used and to what extent in the calculation of each principal component. For example, in the first principal component shown in FIG. 8(a), a rapid population phenomenon is shown in the "T1 period" within the period "from July 1, 2020 to March 31, 2021", and also a significant decrease in population is shown in the "T2 period" within the period "from July 1, 2020 to March 31, 2021". In such an example, based on the principal component loadings of the first principal component, the user can determine that the first principal component is a principal component that strongly reflects the variables corresponding to the "T1 period" and the variables corresponding to the "T2 period". Also, based on such a determination result, the user can consider, for example, how to edit the principal component pedestrian flow data corresponding to the first principal component.
[0076] (Editing Unit 156) The editing unit 156 controls a predetermined pedestrian flow component among the plurality of pedestrian flow components extracted by the inference unit 155 according to the user's purpose (editing operation).
[0077] (Score Change Unit 156a) When information for changing the principal component score corresponding to a predetermined flow component (principal component) among a plurality of flow components is input by the user, the score change unit 156a generates information for changing the principal component score based on the input information. For example, when information for changing the principal component score corresponding to the first principal component is input by the user, the score change unit 156a generates information for changing the current principal component score of the first principal component according to such information.
[0078] Also, the score change unit 156a transmits the generated information to the inference unit 155. As a result, for example, the principal component flow conversion unit 155b generates virtual flow data based on the changed principal component score while maintaining the characteristics of the flow data indicated by the predetermined flow component. For example, the principal component flow conversion unit 155b generates virtual flow data based on the principal component score corresponding to the first principal component, which is the changed principal component score, while maintaining the characteristics of the flow data indicated by the first principal component. The principal component flow conversion unit 155b registers the virtual flow data in the principal component flow storage unit 140a.
[0079] For example, assuming that 16 - variable × 274 pieces (a total of 4384 pieces) of flow data are generated and the clustering unit 153 classifies the flow data into 16 mesh groups. And assuming that function data i (i = 1,..., 16) as shown in FIG. 7 is obtained by the conversion process by the function data conversion unit 154a.
[0080] In such a state, when the user wants to edit the j-th principal component score, which is the principal component score of the j-th principal component of the function data i, according to their own assumptions, the user inputs information indicating the editing content to the information processing apparatus 100. In such a case, the principal component flow transformation unit 155b generates principal component flow data ij based on the j-th principal component score specified by the user among the principal components of the function data i specified by the user and the k-th principal component scores of other function data i (j ≠ k). For example, since the contribution of the flow component to which the principal component score of the function data i is changed changes, and the contributions of other flow components do not change, the flow of the function data i changes only by the change amount of the flow in the flow component for which the principal component score is changed. That is, by changing the j-th principal component score of the function data i, new function data i that is similar to the function data i but differs only in the contribution of the j-th principal component score is generated.
[0081] (Load amount change unit 156b) When information for changing the principal component load amount corresponding to a predetermined flow component (principal component) among a plurality of flow components is input by the user, the load amount change unit 156b generates information for changing the principal component load amount based on the input information. For example, when information for changing the principal component load amount corresponding to the first principal component is input by the user, the load amount change unit 156b generates information for changing the current principal component load amount of the first principal component according to such information.
[0082] Further, the load amount change unit 156b transmits the generated information to the inference unit 155. As a result, for example, the principal component flow transformation unit 155b changes the characteristics of the flow data indicated by a predetermined flow component based on the changed principal component load amount. For example, the principal component flow transformation unit 155b changes the characteristics of the flow data indicated by the first principal component based on the changed principal component load amount. The principal component flow transformation unit 155b registers the flow data with changed characteristics in the principal component flow storage unit 140a.
[0083] (Generation unit 157) The generation unit 157 generates flow data to be provided to the user based on a plurality of flow components including the changed flow components.
[0084] (Principal component pedestrian flow correction unit 157a) When change information for changing the shape of the graph indicated by a predetermined pedestrian flow component (principal component) among a plurality of pedestrian flow components is input by the user, the principal component pedestrian flow correction unit 157a corrects the characteristics of the data indicated by the predetermined pedestrian flow component among the plurality of pedestrian flow components based on the change information. This will be described using the example of FIG. 4(b).
[0085] For example, assume that the user wants to delete the characteristics of population decrease during the "T1 period" and the trend of population decrease during the "T2 period" from the first principal component in which these characteristics are expressed. In such a case, the user can refer to the graph (first principal component pedestrian flow data) indicated by the first principal component extracted from the pedestrian flow data DA11 and perform an operation to delete the characteristics of the "T1 period" and the "T2 period" among the characteristics of the first principal component on the information processing apparatus 100.
[0086] When change information for changing a part of the first principal component pedestrian flow data is input, the principal component pedestrian flow correction unit 157a executes a correction process of deleting the characteristic part corresponding to the input change information among the characteristics included in the first principal component pedestrian flow data. For example, the principal component pedestrian flow correction unit 157a deletes the characteristic part corresponding to the input change information among the characteristics included in the first principal component pedestrian flow data, and executes a process of supplementing the deleted characteristic part by linear interpolation or curve interpolation.
[0087] (Reconstruction unit 157b) The reconstruction unit 157b reconstructs the pedestrian flow data provided to the user based on a plurality of pedestrian flow components including the principal component after the characteristics are controlled. For example, when a part of the characteristics included in the first principal component pedestrian flow data is corrected by the principal component pedestrian flow correction unit 157a, the reconstruction unit 157b adds the pedestrian flow data of the first principal component after the characteristics are corrected and the pedestrian flow data of other principal components (for example, the second principal component to the fourth principal component) to reconstruct the pedestrian flow data provided to the user.
[0088] Among the processing units described so far, the editing unit 156 and the generation unit 157 are processing units corresponding to the crowd flow control unit.
[0089] 〔7. Operation Procedure Example of Information Processing Apparatus〕 Next, an operation procedure example of the information processing apparatus 100 will be described with reference to FIG. 9. FIG. 9 is a flowchart showing the procedure of information processing according to the embodiment.
[0090] When the preprocessing unit 151 acquires raw data of crowd flow (time-series number-of-people data), it performs preprocessing of reorganizing the acquired data by dividing it using statistical methods (e.g., kshape) or qualitative criteria (e.g., a predetermined range such as a mesh) (step S901).
[0091] The function data conversion unit 154a converts the preprocessed crowd flow data into function data using basis functions. For example, the function data conversion unit 154a calculates the basis number m of the function data based on the initial value of function data conversion (step S902).
[0092] Here, among the preprocessed crowd flow data, n time-series data Y at time t are set as the realized values of the function u(t) represented by m bases (step S903). In this case, between the i-th data x_i(t) of the variable x among the n time-series data Y at time t and the i-th data u_i(t) of the variable u among the n time-series data Y at time t, x_i(t)=u_i(t)+ε (ε is a constant) holds.
[0093] And there are n samples of a set of real number data Y represented by m variables. The i-th of the data Y is taken as the data SM in the form of a matrix (n×1), and the (n×m) weight coefficients are taken as the transformation matrix W. The transformation matrix W can also be denoted as the coefficient matrix W or the weight matrix W.
[0094] The PCA learning unit 154b acquires the threshold of the cumulative contribution rate and calculates the number of principal components p for which the cumulative contribution rate is equal to or greater than the threshold from the sample data SM (step S904).
[0095] Then, the PCA learning unit 154b obtains p (p < m) principal components from the transformation matrix W (step S905). Specifically, the PCA learning unit 154b obtains, by unsupervised learning, a transformation matrix W that maximizes the variance of the data in the data SM. For example, the PCA learning unit 154b can obtain, as the transformation matrix W, an eigenvector B corresponding to the maximum eigenvalue. The PCA learning unit 154b obtains, by unsupervised learning, a weight W that maximizes the variance of the inner product of the function data X and the weight W. For example, the PCA learning unit 154b can obtain the weight W as an eigenvalue problem of the covariance function for the function data X. The eigenvector B can be represented by a (p × m) matrix. Also, the principal component score calculation unit 155a can calculate a principal component score S, which is represented by an (n × p) matrix.
[0096] In such a state, the inference unit 155 determines whether to change the input of the inference (step S906). For example, the inference unit 155 can determine to change the input of the inference when user information for changing the principal component score by the score change unit 156a is received, or when user information for changing the principal component loading by the load amount change unit 156b is received. On the other hand, the inference unit 155 may determine not to change the input of the inference when these pieces of user information are not received.
[0097] When it is determined that the input of the inference is not to be changed (step S906; Yes), the principal component flow transformation unit 155b reconstructs a graph indicated by each principal component, that is, a function D(q), based on the flow data for each principal component (step S907).
[0098] On the other hand, when it is determined that the input of the inference is to be changed (step S906; No), the principal component flow transformation unit 155b controls the characteristics of the corresponding principal component flow data based on the input user information (step S908).
[0099] For example, when the principal component flow rate conversion unit 155b receives user information for changing the principal component score, it generates virtual flow rate data based on the principal component score changed according to the user information. For example, the principal component flow rate conversion unit 155b generates virtual flow rate data based on the changed principal component score corresponding to the principal component while maintaining the characteristics of the flow rate data indicated by the principal component specified by the user information.
[0100] As another example, when the principal component flow rate conversion unit 155b receives user information for changing the principal component loading, it changes the characteristics of the flow rate data indicated by the principal component specified by the user information based on the principal component loading changed according to the user information.
[0101] Also, when the process proceeds to step S908, it shifts to step S907. Specifically, the principal component flow rate conversion unit 155b reconstructs the principal component flow rate data controlled according to the user information (user's purpose) and the principal component flow rate data not controlled (outside the user's purpose) respectively.
[0102] Subsequently, the generation unit 157 determines whether to correct the shape of the graph indicated by a predetermined principal component among a plurality of flow rate components (principal components) according to the user information (step S909). For example, when change information for changing the shape of the graph indicated by a predetermined principal component among a plurality of flow rate components is input, the generation unit 157 can determine to correct. More specifically, the generation unit 157 changes the shape of the graph of the principal component flow rate data specified by the user for the principal component flow rate data obtained by decomposing the original flow rate data for each principal component. On the other hand, when change information for changing the shape of the graph indicated by a predetermined principal component among a plurality of flow rate components is not input, the generation unit 157 may determine not to correct.
[0103] When it is determined not to correct (step S909; Yes), the reconstruction unit 157b adds up the flow rate data for each principal component in the uncorrected state to reconstruct one flow rate data (step S910).
[0104] On the other hand, when it is determined that correction is necessary (step S909; No), the principal component passenger flow correction unit 157a determines, based on user information, whether to perform correction that does not retain the shape of the graph indicated by any one of the plurality of passenger flow components (for example, the first principal component) (that is, whether to correct the entire graph or only a part of the graph) (step S911). The principal component passenger flow correction unit 157a directly edits the principal component passenger flow data (unit: person) generated based on the principal component function and the principal component score. Therefore, since the principal component passenger flow data generated from other principal component functions does not change, as a result, other features are retained.
[0105] For example, when the principal component passenger flow correction unit 157a determines that correction is to be performed without retaining the shape of the graph indicated by the first principal component (step S911; Yes), the principal component passenger flow correction unit 157a executes a correction process of deleting the feature portion corresponding to the input change information among the features included in the first principal component passenger flow data. For example, the principal component passenger flow correction unit 157a deletes the feature portion corresponding to the input user information among the features included in the first principal component passenger flow data, and executes a process of supplementing the deleted feature portion by linear interpolation or curve interpolation (step S912).
[0106] On the other hand, when the principal component passenger flow correction unit 157a determines that correction is to be performed while retaining the shape of the graph (step S911; No), the principal component passenger flow correction unit 157a expands the same period as the control point on the principal component passenger flow data to the control point (for example, a holiday or a quantile), and for the periods other than the control point, expands them by determining, for example, by linearly interpolating the expansion rate from the rate of the control point (step S913).
[0107] Also, when the process proceeds to step S912 or S913, the process proceeds to step S910. Specifically, the reconstruction unit 157b adds the principal component passenger flow data corrected by linear interpolation or the like and the uncorrected principal component passenger flow data to reconstruct one passenger flow data (step S910).
[0108] So far, the operation procedure example of the information processing apparatus 100 has been described. According to the example in FIG. 9, steps S901 to S903 are the preprocessing of the pedestrian flow data. Also, steps S904 to S905 are the principal component analysis of the pedestrian flow data. Further, steps S906 to S908 are the processes of editing the characteristics of the pedestrian flow data according to the input on the user (for example, administrator) side. On the other hand, steps S909 to S913 are the processes of editing the characteristics of the pedestrian flow data according to the input on the user (for example, service user) side.
[0109] Here, a specific example of the generation method of the principal component pedestrian flow data ij is shown using the example in FIG. 9. First, time series data is approximated by m functions g_1 to g_m as functions. x_i(t)=u_i(t)+ε (ε is a constant) holds. The estimated coefficients x_i_1 to x_i_m are grouped together to define an n×m design matrix D. When the design matrix D is subjected to principal component analysis, an eigenvector B=(b_1,···,b_p) is obtained.
[0110] The j-th principal component function is obtained by the inner product of the eigenvector b_j and the function vector (g_1, g_2,···,g_m)'. Therefore, m principal component functions (h_1, h_2,···,h_m) are obtained by the inner product of the transpose of the eigenvector B and the function vector. Also, the principal component score vector (s_i_1,···,s_i_m) is obtained by the inner product of the coefficient vector (x_i_1,···,x_i_m) and the eigenvector B.
[0111] Time series_i can be represented as a linear combination of m principal component functions (h_1, h_2,···,h_m), and the coefficients at this time are the principal component scores. Specifically, in Time series_i = s_i_1×h_1 +···+ s_i_m×h_m + ξ, s_i_j×h_j is a function obtained by multiplying the function h_j by a constant and is not time series data. When x = 0,···,273 is substituted, it returns to the time series data, which becomes the principal component pedestrian flow data ij.
[0112] 〔8. Hardware Configuration〕 The information processing device 100 according to the above embodiment is realized by a computer 1000 having a configuration as shown in Fig. 10. Fig. 10 is a hardware configuration diagram showing an example of the computer 1000 that realizes the functions of the information processing device 100. The computer 1000 has a CPU 1100, a RAM 1200, a ROM 1300, an HDD 1400, a communication interface (I / F) 1500, an input / output interface (I / F) 1600, and a media interface (I / F) 1700.
[0113] The CPU 1100 operates and controls each unit based on programs stored in the ROM 1300 or the HDD 1400. The ROM 1300 stores a boot program executed by the CPU 1100 when the computer 1000 starts up, programs that depend on the hardware of the computer 1000, and the like.
[0114] The HDD 1400 stores programs executed by the CPU 1100, data used by the programs, etc. The communication interface 1500 receives data from other devices via the communication network 50 and sends it to the CPU 1100, and transmits data generated by the CPU 1100 to other devices via the communication network 50.
[0115] The CPU 1100 controls output devices such as a display and a printer, and input devices such as a keyboard and a mouse, via the input / output interface 1600. The CPU 1100 acquires data from the input devices via the input / output interface 1600. The CPU 1100 also outputs generated data to the output devices via the input / output interface 1600.
[0116] The media interface 1700 reads the programs or data stored in the recording medium 1800 and provides them to the CPU 1100 via the RAM 1200. The CPU 1100 loads such a program from the recording medium 1800 onto the RAM 1200 via the media interface 1700 and executes the loaded program. The recording medium 1800 is, for example, an optical recording medium such as a DVD (Digital Versatile Disc), a PD (Phase change rewritable Disk), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory, etc.
[0117] For example, when the computer 1000 functions as the information processing apparatus 100 according to the embodiment, the CPU 1100 of the computer 1000 realizes the functions of each processing unit by executing a program (for example, the information processing program according to the embodiment) loaded onto the RAM 1200. Also, data in the storage unit is stored in the HDD 1400. The CPU 1100 of the computer 1000 reads and executes these programs from the recording medium 1800, but as another example, these programs may be acquired from other devices via the communication network 50.
[0118] [9. Others] Also, each component of each illustrated device is conceptually functional and does not necessarily have to be physically configured as shown in the figure. That is, the specific form of the distribution and integration of each device is not limited to that shown in the figure, and all or part of it can be functionally or physically distributed and integrated in any unit according to various loads, usage situations, etc.
[0119] [10. Summary] As described above, the embodiments of the present application have been described in detail based on several drawings, but these are examples, and the present invention can be implemented in other forms with various modifications and improvements based on the knowledge of those skilled in the art, starting from the aspects described in the column of the disclosure of the invention. [Explanation of Reference Numerals]
[0120] 1 System 10 User Device 100 Information Processing Device 111 External Interface Unit 112 Management Interface Unit 113 Input / Output Interface Unit 120 Input Storage Unit 130 Memory Unit 140 Output Storage Unit 151 Preprocessing Unit 152 Operation Unit 153 Clustering Unit 154 Learning Unit 155 Inference Unit 156 Editing Unit 157 Generation Unit
Claims
1. A conversion unit that converts the flow data of people, which is time-series flow data of the population in a predetermined range, into function data, An extraction unit that extracts a plurality of flow components of people by performing principal component analysis on the function data, A flow control unit that controls, among the plurality of flow components of people, a predetermined flow component of people to be changed according to the purpose of the user, A generation unit that generates flow data of people to be provided to the user based on the plurality of flow components of people including the changed flow component of people, An information processing apparatus comprising the same.
2. The flow data of people is multidimensional data including a plurality of variables for the observation of the population in a predetermined period for each predetermined range, The conversion unit converts each of the flow data of people that exists according to the number of the variables into function data, The information processing apparatus according to Claim 1.
3. The extraction unit calculates, for each of the flow components of people, a principal component score corresponding to the plurality of variables and a principal component loading amount indicating the relationship between each variable among the plurality of variables when extracting the plurality of flow components of people, The information processing apparatus according to Claim 2.
4. When information for changing the principal component score corresponding to a predetermined flow component of people among the plurality of flow components of people is input by the user, the flow control unit controls the characteristics of the data indicated by the predetermined flow component of people based on the changed principal component score The information processing apparatus according to Claim 3.
5. When information for changing the principal component loading amount corresponding to a predetermined flow component of people among the plurality of flow components of people is input by the user, the flow control unit controls the characteristics of the data indicated by the predetermined flow component of people based on the changed principal component loading amount The information processing apparatus according to Claim 3.
6. When change information for changing the shape of a graph indicated by a predetermined flow component of people among the plurality of flow components of people is input by the user, the flow control unit controls the characteristics of the data indicated by the predetermined flow component of people among the plurality of flow components of people based on the change information The information processing apparatus according to Claim 1.
7. The generation unit generates flow data of people to be provided to the user based on the plurality of flow components of people including the predetermined flow component of people whose characteristics have been controlled The information processing apparatus according to any one of Claims 4 to 6.
8. An information processing method executed by an information processing apparatus, A conversion step of converting the flow data of people, which is time-series flow data of the population in a predetermined range, into function data, An extraction step of extracting a plurality of people flow components by performing principal component analysis on the function data; A people flow control step of controlling so that a predetermined people flow component among the plurality of people flow components is changed according to the purpose of the user; A generation step of generating people flow data to be provided to the user based on the plurality of people flow components including the changed people flow component; An information processing method including the above.
9. A conversion procedure for converting time-series people flow data, which is people flow data of the population in a predetermined range, into function data; An extraction procedure for extracting a plurality of people flow components by performing principal component analysis on the function data; A people flow control procedure for controlling so that a predetermined people flow component among the plurality of people flow components is changed according to the purpose of the user; A generation procedure for generating people flow data to be provided to the user based on the plurality of people flow components including the changed people flow component; An information processing program for causing a computer to execute the above.
Citation Information
Patent Citations
Construction facility control system and program
JP2009294887A
Image processing apparatus, image processing method, and program
JP2019067208A
People flow analysis program, people flow analysis method, and people flow analysis system
JP7027605B1