Machine learning method, computer-readable recording medium, and information processing device

WO2026204271A1PCT designated stage Publication Date: 2026-10-01FUJITSU LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/008802
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2026-03-06
Publication Date
2026-10-01

Smart Images

  • Figure JP2026008802_01102026_PF_FP_ABST
    Figure JP2026008802_01102026_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device generates, on the basis of an initial feature and a first feature, a third feature to be input to a classifier that performs walking recognition. The initial feature is generated on the basis of a first video of normal walking in a data set including point cloud videos of a walking recognition target, and the first feature is extracted from a second video in the data set, the second video including the normal walking and walking with an external attribute.
Need to check novelty before this filing date? Find Prior Art

Description

Machine learning method, computer-readable recording medium, and information processing apparatus

[0001] The present invention relates to a machine learning method, a computer-readable recording medium, and an information processing apparatus.

[0002] Gait recognition is a technology for performing individual analysis and identification based on a person's gait pattern, and can be applied in various fields such as security, healthcare, and sports.

[0003] Conventionally, technologies for performing gait recognition using point cloud data are known.

[0004] Chuanfu Shen, Fan Chao, Wei Wu, Rui Wang, George Q. Huang, Shiqi Yu"LidarGait: Benchmarking 3D Gait Recognition with Point Clouds", [online], [retrieved November 19, 2024], Internet <https: / / ieeexplore.ieee.org / document / 10205455> Xiang Li, Yasushi Makihara, Chi Xu, Yasushi Yagi, Mingwu Ren"Gait Recognition via Semi-supervised Disentangled Representation Learning to Identity and Covariate Features", [online], [retrieved November 19, 2024], Internet <https: / / ieeexplore.ieee.org / document / <9156701> Chao Fan, Yunjie Peng, Chunshui Cao, Xu Liu, Saihui Hou, Jiannan Chi"GaitPart: Temporal Part-Based Model for Gait Recognition", [online], [retrieved January 29, 2025], Internet <https: / / ieeexplore.ieee.org / document / <9156784>

[0005] ​However, with the above technology, for example, even if a pedestrian is carrying an item (i.e., if external attributes such as the item are included), the entire pedestrian's body, including the item (external attribute), is treated as a single processing target, which sometimes resulted in a decrease in the performance of the pedestrian recognition process.

[0006] Conventional gait recognition technologies using point cloud data often treat the entire body, including external attributes, as a monolithic object through the projection of the point cloud, or divide it into multiple parts to be processed as multiple objects. Because they cannot consider the existence of external attributes, the recognition process can become complicated.

[0007] One aspect of this is the aim to enable highly accurate gait recognition.

[0008] In the first proposal, the machine learning method involves a computer generating a third set of features to be input into a classifier that performs gait recognition. This is done by using an initial set of features generated from a first video of normal walking, and a second set of features extracted from a second video that includes both normal walking and walking with external attributes, from a dataset containing point cloud videos of the target of gait recognition.

[0009] According to one embodiment, it is possible to achieve highly accurate gait recognition.

[0010] Figure 1 is a diagram illustrating the information processing device according to Example 1. Figure 2 is a diagram illustrating the feature extraction block according to Example 1. Figure 3 is a diagram illustrating the initial feature extraction block according to Example 1. Figure 4 is a diagram illustrating the initial feature extraction block according to Example 1. Figure 5 is a functional block diagram showing the functional configuration of the information processing device according to Example 1. Figure 6 is a flowchart illustrating the flow of the training phase according to Example 1. Figure 7 is a flowchart illustrating the flow of the training phase according to Example 1. Figure 8 is a flowchart illustrating the flow of the application phase according to Example 1. Figure 9 is a diagram illustrating an example of hardware configuration.

[0011] The following describes in detail, with reference to the drawings, embodiments of the machine learning method, computer-readable recording medium, and information processing apparatus disclosed herein. However, these embodiments do not limit the present invention. Furthermore, each embodiment can be combined as appropriate within a consistent scope.

[0012] [1. Introduction] Gait recognition is a technology that analyzes and identifies individuals based on their walking patterns, and can be applied in various fields such as security, healthcare, and sports.

[0013] In the field of security, it could be applied, for example, as a pedestrian recognition system for analyzing and identifying individuals in public spaces. For instance, it could detect abnormal behavior or potential threats based on walking patterns. It could also be applied to preventative measures and automated check-ins.

[0014] In the healthcare field, this technology can be applied to areas such as disease diagnosis, monitoring, and rehabilitation. For example, analyzing gait patterns can help diagnose symptoms of Parkinson's disease, stroke, and other neurological conditions. Furthermore, monitoring and analyzing a patient's gait over a period of time can enable the development of personalized rehabilitation programs.

[0015] By analyzing changes in gait patterns, it may be possible to detect early symptoms of neurological diseases such as Parkinson's disease and Alzheimer's disease. It can also provide valuable feedback on the recovery status of patients recovering from stroke, surgery, or injury. Furthermore, it can be applied to the design and customization of prosthetics and corrective devices to achieve high functionality tailored to the patient.

[0016] In the field of sports, it can be applied to, for example, overall diagnosis and performance analysis. For instance, analyzing an athlete's gait can lead to improved performance, reduced injury risk, and optimized training plans. It can also be applied to understanding biomechanics and exercise mechanics.

[0017] Point cloud data has advantages in terms of privacy protection because it represents individuals as a collection of points rather than as an image. Furthermore, point cloud data enables spatial representation, making it applicable to a variety of real-world situations. Therefore, it enables pedestrian recognition in diverse scenarios. For example, pedestrian recognition is possible regardless of the location, the items the pedestrian is holding, or the actions they are taking.

[0018] Conventionally, technologies for gait recognition using point cloud data are known. However, with these technologies, for example, even when a pedestrian is carrying an item (i.e., when external attributes such as carried items are included), the entire pedestrian's body, including the carried item (external attribute), is treated as a single processing target, which can sometimes lead to a decrease in the performance of the gait recognition process. Furthermore, these technologies sometimes rely heavily on neural networks, which can lead to the overlooking of specific external attributes.

[0019] Conventional gait recognition technologies using point cloud data often treat the entire body, including external attributes, as a monolithic processing target through point cloud projection, or divide it into multiple parts to be processed as multiple targets. Because they cannot consider the existence of external attributes, the recognition process can become complex.

[0020] For example, in the technology disclosed in Non-Patent Document 1, the entire body of a pedestrian, including external attributes, is treated as a single processing target. Therefore, if external attributes are included, the performance of the pedestrian recognition processing may decrease.

[0021] Furthermore, when inputting gait data with external attributes into a model for gait recognition, the model often focuses more on the external attributes than on the pedestrian. Also, even when the entire body is divided into multiple parts, it is often impossible to capture the actual gait pattern or remove the external attributes, resulting in a high computational load.

[0022] Furthermore, as an example of gait recognition technology other than point cloud data, a technology that uses silhouette image data for gait recognition, such as the technology disclosed in Non-Patent Document 3, is known. However, in such technologies, changes in lighting, camera angle, and surface type can affect the accuracy of gait recognition. For example, accuracy decreases when clothing changes. Also, because features at the sub-unit level are processed in time series, the computational load becomes high.

[0023] Many conventional gait recognition technologies process binary silhouettes rather than point cloud data. However, gait recognition using binary silhouettes cannot determine which parts are the front and back of the body, potentially compromising the information necessary for accurate gait recognition. While using depth image data can prevent the loss of information necessary for accurate gait recognition, it becomes more susceptible to the influence of external attributes present in the data.

[0024] In this respect, point cloud data, unlike image data such as silhouette image data and depth image data, can achieve an appropriate 3D representation of the real world.

[0025] This application was made in view of the above, and aims to enable highly accurate gait recognition regardless of the presence or absence of external attributes.

[0026] As will be explained in detail later, the following embodiment uses four blocks: initial feature extraction block BL1, feature extraction block BL2, problem recognition feature enhancement block BL3, and classification block BL4. The output results of the two blocks, initial feature extraction block BL1 and feature extraction block BL2, are used to perform problem recognition feature enhancement block BL3. Then, the output results of problem recognition feature enhancement block BL3 are used to perform classification block BL4.

[0027] In the embodiment described below, unlike the above technology, instead of performing gait recognition as a monolithic object including external attributes, latent features (hereinafter referred to as "initial features") are generated from a gait video of normal walking without external attributes (hereinafter referred to as "first video"), and mutual attention is performed between the features (hereinafter referred to as "first features") extracted from a gait video including both normal walking and walking with external attributes (hereinafter referred to as "second video") and the initial features, thereby enabling highly accurate gait recognition regardless of the presence or absence of external attributes.

[0028] To make this possible, the initial feature extraction block BL1 first extracts features (hereinafter referred to as "secondary features" as appropriate) from the first video. A predetermined number of frames (point cloud frame count) is selected, and secondary features are extracted from each frame (point cloud frame) of the first video. This initial step makes it possible to extract features that represent the walking pattern of normal walking. Then, the secondary features are aggregated using spatiotemporal attention to generate the initial features. These initial features are comprehensive vectors.

[0029] Then, in the feature extraction block BL2, the first feature is extracted from the second video using the same feature extractor (corresponding to feature extractor MO1) that was used to extract the second feature in the initial feature extraction block BL1. This feature extractor is a deep neural network (DNN).

[0030] Then, in the problem recognition feature enhancement block BL3, mutual attention is performed between the first feature and the initial feature. Mutual attention allows for minimizing the influence of external attributes by focusing on the relevant part of the body (the part including external attributes). In other words, it becomes possible to filter out external attributes such as umbrellas and bags with high accuracy. That is, by focusing on the relevant part of the body through mutual attention, features related to external attributes can be filtered with high accuracy, making it possible to generate features (corresponding to the third feature described later) with the influence of external attributes minimized.

[0031] By incorporating the initial feature extraction block BL1 and generating initial features, it becomes possible to distinguish between normal walking patterns. This eliminates the need to add annotations, reducing overhead. Furthermore, even if the point cloud data contains unknown external attributes, it is possible to effectively train the pedestrian's walking pattern, enabling highly accurate walking recognition (and the provision of a walking recognition system).

[0032] In the following embodiment, we will explain using the example where the object of gait recognition (the object of gait recognition) is a pedestrian (i.e., a person), but it is not limited to people; for example, it could be an animal such as a dog or cat, or even a robot.

[0033] [2. Explanation of Information Processing] Next, the information processing according to the embodiment will be explained. Figure 1 is a diagram illustrating the information processing device 10 according to Embodiment 1. The information processing device 10 shown in Figure 1 is an example of a computer device that enables highly accurate gait recognition regardless of the presence or absence of external attributes. As shown in Figure 1, the information processing device 10 includes an initial feature extraction block BL1, a feature extraction block BL2, a task recognition feature enhancement block BL3, and a classification block BL4.

[0034] (Details of Feature Extraction Block BL2) First, the details of Feature Extraction Block BL2 will be explained using Figure 2. Figure 2 is a diagram illustrating Feature Extraction Block BL2 according to Embodiment 1. In Feature Extraction Block BL2, the information processing device 10 takes video MV1 (corresponding to the second video), which is a point cloud video (walking sequence data) of a pedestrian (walking recognition target), as input and performs preprocessing (corresponding to Module MD1). In Module MD1, the information processing device 10 resizes each point cloud frame (CF1 to CFn) of video MV1 so that all of them become predetermined points, and adjusts them to a predetermined number of frames by trimming and expanding (for example, by trimming the long parts and expanding the short parts). In Module MD1, the information processing device 10 may also perform processing such as averaging and centralization, and may perform it at any timing.

[0035] For the sake of explanation, the following example will use the case where each point cloud frame (CF1 to CFn) in video MV1 is resized to 1024 points, and the number of frames is adjusted to 16. However, the explanation is not limited to this example. Note that if the number of frames is reduced to 5 or 6 frames, the accuracy of gait recognition will decrease, while if the number of frames is increased to 100 or 200 frames, the accuracy of gait recognition will increase, but the computational load will also increase, making it less user-friendly. The number of frames, 16, is set based on a predetermined frame rate that is assumed to be able to capture a sufficient amount of data from the perspective of gait recognition accuracy, but the explanation is not limited to this example.

[0036] The information processing device 10 inputs the pre-processed video MV1 (corresponding to video AMV1) from module MD1 into the feature extractor MO1. In other words, the feature extractor MO1 receives each of the 16 adjusted point cloud frames (ACF1 to ACFn) (since there are 16 frames, ACFn is ACF16). As a result, a first feature is extracted for each frame.

[0037] (Details of Initial Feature Extraction Block BL1) Next, the details of Initial Feature Extraction Block BL1 will be explained using Figures 3 and 4. Figures 3 and 4 are diagrams illustrating Initial Feature Extraction Block BL1 according to Example 1. In Initial Feature Extraction Block BL1, initial features are generated in the training phase. In the application phase (test phase), the most recent initial feature generated in the training phase is used.

[0038] Similar to the case in Figure 2, the information processing device 10 performs preprocessing on the pedestrian point cloud video as input. In Figure 3, preprocessing (corresponding to module MD11) is performed on the video MV11 (corresponding to the first video) as input. In module MD11, the information processing device 10 resizes each point cloud frame (CF11 to CF1n) of video MV11 so that each frame has 1024 points, and adjusts it to 16 frames by trimming and expanding. This is done for each first video (videos MV11 to MV1N).

[0039] The information processing device 10 inputs the pre-processed video MV11 (corresponding to video AMV11) from module MD11 into the feature extractor MO1. In other words, the feature extractor MO1 receives each adjusted point cloud frame (ACF11 to ACF1n) in 16 frames (since there are 16 frames, ACFn is ACF116). As a result, a second feature is extracted for each frame. This is done for each first video (video MV11 to MV1N). For example, in the case of video MV12, each adjusted point cloud frame (ACF21 to ACF2n) is input in 16 frames (since there are 16 frames, ACF2n is ACF216).

[0040] In this manner, the information processing device 10 inputs each of the 16 frames of the first video of normal walking into the feature extractor MO1. The information processing device 10 aggregates the second features extracted frame by frame from each first video via the feature extractor MO1 into the initial features using spatiotemporal attention (corresponding to module MD2). Furthermore, the information processing device 10 recalculates the initial features after a predetermined number of training iterations to update the latest initial features.

[0041] During the training phase, the initial features are updated after a predetermined number of training iterations. Upon transitioning to the application phase, the most recent initial features are used for consistency and reliability of the evaluation.

[0042] For the sake of explanation, the following example will use a predetermined number of training iterations of 100 and the initial features are recalculated every 100 training iterations, but the explanation is not limited to this example.

[0043] (Details of the problem recognition feature enhancement block BL3) Returning to the explanation of Figure 1, the information processing device 10 obtains a first feature extracted for each frame by the feature extraction block BL2 and an initial feature generated based on the second feature extracted for each frame by the initial feature extraction block BL1.

[0044] In the task recognition feature emphasis block BL3, the information processing apparatus 10 performs mutual attention between the first feature and the initial feature (corresponding to the module MD3). By performing mutual attention, while features essential for gait recognition are emphasized, it becomes possible to suppress features related to unnecessary external attributes and occlusion. In the training phase, the result of the mutual attention is trained. In the application phase, necessary features indispensable for accurate gait recognition among the first features are identified. Further, by filtering features related to external attributes with high precision, it becomes possible to generate a feature (corresponding to a third feature described later) in which the influence of external attributes is minimized.

[0045] Furthermore, in the task recognition feature emphasis block BL3, the information processing apparatus 10 performs temporal pooling (corresponding to the module MD4). The information processing apparatus 10 aggregates the features emphasized by the module MD3 for each frame in order to capture temporal dynamics. Hereinafter, this aggregated feature is referred to as the third feature as appropriate. As a result, the third feature with high identification accuracy and high robustness is generated.

[0046] (Details of classification block BL4) In the classification block BL4, the information processing apparatus 10 performs classification (corresponding to the module MD5). The information processing apparatus 10 performs final identification prediction or classification prediction of the video MV1 (for example, predicts the identification ID of a pedestrian) by using a fully connected layer and a SoftMax classifier on the third feature. Note that the use of the SoftMax classifier is an example, and the present invention is not particularly limited to this example. That is, any classifier may be used without being limited to the SoftMax classifier as long as it is a classifier capable of performing identification prediction or classification prediction with features as input. Further, the information processing apparatus 10 may use a SoftMax classifier trained by triplet loss widely used for detection-based machine learning.

[0047] In the training phase, the information processing apparatus 10 trains the feature extractor MO1 using a loss function until the loss is minimized. The information processing apparatus 10 trains the feature extractor MO1 such that the loss is minimized through iterative processing of a predetermined number of training iterations (for example, 40000 training iterations).

[0048] The details of the initial feature extraction block BL1, feature extraction block BL2, task-aware feature enhancement block BL3, and classification block BL4 have been described above. The above is an overview of the information processing apparatus 10 according to the first embodiment.

[0049] Based on the above, the information processing apparatus 10 can provide a high-precision walking recognition system that is not affected by external attributes present in data. Furthermore, the information processing apparatus 10 can directly process point cloud data without going through intermediate representations obtained by intermediate projection. Furthermore, through the above-described training, the information processing apparatus 10 can accurately determine where to focus in point cloud data, thereby effectively removing unnecessary external attributes.

[0050] Furthermore, by focusing on main walking features based on initial features, it is possible to reduce noise caused by external changes. Furthermore, through mutual attention, dynamic adjustment is performed to enable focus on appropriate walking patterns, which makes it possible to improve robustness. Furthermore, directly processing point cloud data reduces the need for models with a large resource burden, thereby enabling the introduction of the system in a wide range of applications. Furthermore, task-aware feature enhancement improves performance under various conditions, thereby enabling provision of a highly versatile system. Furthermore, processing with point cloud data enhances privacy protection, so the system can also be applied to uses requiring a high security level.

[0051] [3. Configuration of Information Processing Apparatus] Next, the configuration of the information processing apparatus 10 according to the first embodiment will be described. FIG. 5 is a functional block diagram showing the functional configuration of the information processing apparatus 10 according to the first embodiment. As shown in FIG. 5, the information processing apparatus 10 includes a communication unit 11, a storage unit 12, and a control unit 20. Note that the information processing apparatus 10 is not limited to the configuration illustrated, and may include a display unit and the like.

[0052] The communication unit 11 executes communication with other devices. For example, the communication unit 11 receives input data. The communication unit 11 also receives various instructions, data, and the like from an administrator terminal used by an administrator.

[0053] The memory unit 12 stores various data and programs executed by the control unit 20. For example, the memory unit 12 stores walking video DB 13, initial feature DB 14, and machine learning model DB 15.

[0054] The walking video DB13 is a database that stores walking videos from various real-world situations. For example, the walking video DB13 stores training data used to train the feature extractor MO1.

[0055] The initial feature DB14 is a database that stores the most recent initial features. For example, the initial feature DB14 stores the most recent initial features used to generate the third feature in the application phase.

[0056] The Machine Learning Model DB 15 is a database that stores models generated by machine learning. For example, the Machine Learning Model DB 15 stores models using DNNs. For example, the Machine Learning Model DB 15 stores models using neural networks and other machine learning algorithms. Note that the models stored in the Machine Learning Model DB 15 may be generated by other devices.

[0057] The models stored in the machine learning model DB 15 are the various models described in Example 1. For example, the models stored in the machine learning model DB 15 are the feature extractor MO1 and the SoftMax classifier (the SoftMax classifier is just one example; any classifier can be used).

[0058] Feature extractor MO1 is a model used in the initial feature extraction block BL1 and feature extraction block BL2. Feature extractor MO1 takes point cloud data (which can be point cloud video or point cloud frames) as input and outputs either a first or second feature. Examples of feature extractors MO1 include DGCNN and PointNet++.

[0059] The SoftMax classifier is a model used in classification block BL4. The SoftMax classifier takes a third feature as input and outputs an identification prediction result or a classification prediction result (such as a pedestrian identification ID).

[0060] The control unit 20 is the processing unit that controls the entire information processing device 10. For example, the control unit 20 includes an acquisition unit 21, a preprocessing unit 22, a first generation unit 23, an extraction unit 24, a second generation unit 25, a classification unit 26, a loss unit 27, and a training unit 28.

[0061] The acquisition unit 21 acquires a dataset containing point cloud videos of the target of gait recognition. The acquisition unit 21 acquires the latest initial features along with the feature extractor MO1 that was trained in the training phase. The acquisition unit 21 acquires the feature extractor MO1 and initializes the number of training iterations.

[0062] The preprocessing unit 22 performs preprocessing on each point cloud frame of the point cloud video. The preprocessing unit 22 adjusts each point cloud frame of the point cloud video to a predetermined number of frames. That is, the preprocessing unit 22 generates a predetermined number of point cloud frames. The preprocessing unit 22 generates a predetermined number of first point cloud frames from a first video of normal walking, and generates a predetermined number of second point cloud frames from a second video that includes both normal walking and walking with external attributes.

[0063] In this case, the preprocessing unit 22 may perform processing such as averaging and normalization on each point cloud frame of the point cloud video, or it may perform processing such as averaging and normalization on each point cloud frame after it has been adjusted to a predetermined number of frames. In other words, the preprocessing unit 22 may perform processing such as averaging and normalization at any stage during preprocessing, such as before resizing, after resizing, before length adjustment such as trimming or expansion, after length adjustment such as trimming or expansion, before adjusting to a predetermined number of frames, or after adjusting to a predetermined number of frames.

[0064] The first generation unit 23 generates initial features. The first generation unit 23 generates initial features based on the first video. The first generation unit 23 inputs the first point cloud frame into the feature extractor MO1 (which may be a feature extractor MO1 trained by the training unit 28 described later) and generates initial features by aggregating the extracted second features using spatiotemporal attention. The first generation unit 23 generates initial features by aggregating multiple second features extracted from each video of the first video using the feature extractor MO1. The first generation unit 23 generates initial features by aggregating multiple second features extracted from each video of the first video for each first point cloud frame using the feature extractor MO1.

[0065] The extraction unit 24 extracts a first feature. The extraction unit 24 extracts a first feature from the second video. The extraction unit 24 extracts a first feature by inputting the second point cloud frame into the feature extractor MO1 (which may be a feature extractor MO1 trained by the training unit 28 described later). The extraction unit 24 uses the feature extractor MO1 to extract multiple first features from the second video for each second point cloud frame.

[0066] The second generation unit 25 generates a third feature. The second generation unit 25 generates a third feature for input to the SoftMax classifier that performs gait recognition. The second generation unit 25 generates a third feature based on the initial feature and the first feature. The second generation unit 25 generates a third feature based on the initial feature generated by the first generation unit 23 and the first feature extracted by the extraction unit 24.

[0067] The second generation unit 25 performs mutual attention between the initial features and the first features. The second generation unit 25 focuses on the relevant part of the body and performs mutual attention to minimize the influence of external attributes. The second generation unit 25 generates a third feature by aggregating the features that have been emphasized by the mutual attention between the initial features and the first features. The second generation unit 25 generates a third feature by aggregating the features that have been emphasized by the mutual attention between the initial features and the first features in a way that captures temporal dynamics (i.e., by performing time pooling). The second generation unit 25 generates a third feature by aggregating the features for each second point cloud frame that have been emphasized by the mutual attention between the initial features and the first features in a way that captures temporal dynamics.

[0068] The classification unit 26 performs gait recognition. The classification unit 26 performs gait recognition using a third feature. The classification unit 26 performs gait recognition using a third feature generated by the second generation unit 25. The classification unit 26 performs discrimination prediction or classification prediction using a fully connected layer and a SoftMax classifier. The classification unit 26 further refines the feature representation of the third feature using a fully connected layer and inputs it to the SoftMax classifier to perform discrimination prediction or classification prediction. The classification unit 26 performs discrimination prediction or classification prediction using a SoftMax classifier trained with triplet loss.

[0069] The loss unit 27 calculates and determines the loss. The loss unit 27 calculates the loss using a loss function. The loss unit 27 calculates the loss based on the gait recognition result. The loss unit 27 determines whether the loss has been minimized. If the loss unit 27 determines that the loss has been minimized, it determines to terminate the training phase. If the loss unit 27 determines that the loss has not been minimized, it determines to update the number of training iterations, acquire a training batch, and continue the training phase.

[0070] The training unit 28 trains the feature extractor MO1. The training unit 28 trains the feature extractor MO1 based on the gait recognition results. The training unit 28 trains the feature extractor MO1 based on the gait recognition results from the classification unit 26. The training unit 28 trains the feature extractor MO1 based on the error between the gait recognition results from the classification unit 26 and the correct data of the gait recognition target. The training unit 28 trains the feature extractor MO1 based on the loss determination result calculated by the loss unit 27 based on the gait recognition results from the classification unit 26. The training unit 28 trains the feature extractor MO1 until the loss calculated by the loss unit 27 is minimized.

[0071] The training unit 28 trains the same feature extractor MO1, which is used both when the first generation unit 23 generates initial features based on the first video and when the extraction unit 24 extracts the first features from the second video. The training unit 28 updates the initial features each time the number of training iterations for training the feature extractor MO1 reaches a predetermined threshold, and trains the feature extractor MO1 using the latest updated initial features. The training unit 28 updates the number of iterations and obtains a training batch. After updating the initial features, the training unit 28 updates the number of iterations and obtains a training batch. If the loss unit 27 determines that the loss has not been minimized, the training unit 28 updates the number of iterations and obtains a training batch.

[0072] The following describes the overview of the information processing device 10 according to Example 1, divided into a training phase and an application phase, using a flowchart.

[0073] (Training Phase Flow 1) Figure 6 is a flowchart showing the training phase flow according to Example 1. The acquisition unit 21 acquires a dataset including 3D point cloud videos of pedestrians (S11). The preprocessing unit 22 performs preprocessing, resizing each point cloud frame to 1024 points, and adjusting it to 16 frames by trimming and expanding (S12).

[0074] The first generation unit 23 generates initial features in the initial feature extraction block BL1 (S13). Specifically, the first generation unit 23 inputs each of the 16 frames of the first video of normal walking into the feature extractor MO1 and aggregates the extracted second features into the initial features using spatiotemporal attention. The first generation unit 23 updates the initial features every 100 training iterations.

[0075] The extraction unit 24 extracts the first feature in the feature extraction block BL2 (S14). Specifically, the extraction unit 24 inputs the second video, which includes walking with external attributes for 16 frames, into the feature extractor MO1 and extracts the first feature for each frame.

[0076] The second generation unit 25 performs mutual attention between the first feature and the initial feature in the task recognition feature enhancement block BL3 (S15). The second generation unit 25 also performs time pooling and aggregates the features of each frame to generate the third feature (S16).

[0077] The classification unit 26 performs identification or classification prediction using the third feature in the classification block BL4 (S17). The loss unit 27 calculates the loss using the loss function (S18). The training unit 28 then repeatedly trains the feature extractor MO1 until the loss is minimized (S19).

[0078] (Training Phase Flow 2) Figure 7 is a more detailed flowchart showing the flow of the training phase. A brief explanation follows.

[0079] The preprocessing unit 22 performs preprocessing of the training data (which may include mean centralization and normalization) (S101). The acquisition unit 21 acquires the feature extractor MO1 and initializes the number of training iterations (S102). The training unit 28 updates the number of training iterations and acquires a training batch (S103).

[0080] The training unit 28 determines whether the number of training iterations is divisible by 100 (S104). If the training unit 28 determines that the number of training iterations is divisible by 100 (S104: Yes), it determines to sample normal walking data (S105). The first generation unit 23 extracts the second feature (S106). The first generation unit 23 generates the initial feature and updates the latest initial feature (S107). Then it returns to S103 and processes again.

[0081] On the other hand, if the training unit 28 determines that the number of training iterations is not divisible by 100 (S104: No), it decides to extract the first feature (S108). The acquisition unit 21 acquires the latest initial feature (S109). The second generation unit 25 performs mutual attention (S110). The training unit 28 trains the feature extractor MO1 (S111). The loss unit 27 calculates the loss (S112).

[0082] The training unit 28 determines whether the loss has been minimized (S113). If the training unit 28 determines that the loss has been minimized (S113: Yes), it terminates the information processing (i.e., terminates the training phase). On the other hand, if the training unit 28 determines that the loss has not been minimized (S113: No), it returns to S103 and processes again.

[0083] (Flow of the Application Phase) Figure 8 is a flowchart showing the flow of the application phase according to Example 1. The acquisition unit 21 acquires a dataset including 3D point cloud video of pedestrians (S21). The preprocessing unit 22 performs preprocessing, resizing each point cloud frame to 1024 points, and adjusting it to 16 frames by trimming and expanding (S22).

[0084] The acquisition unit 21 acquires the latest initial features in the initial feature extraction block BL1, along with the feature extractor MO1 which was trained in the training phase (S23).

[0085] The extraction unit 24 extracts the first feature in the feature extraction block BL2 using the feature extractor MO1 that was trained in the training phase (S24). Specifically, the extraction unit 24 inputs the second video, which includes walking with external attributes for 16 frames, into the feature extractor MO1 that was trained in the training phase, and extracts the first feature for each frame.

[0086] The second generation unit 25 performs mutual attention between the first feature and the initial feature in the task recognition feature enhancement block BL3 (S25). The second generation unit 25 also performs time pooling and aggregates the features of each frame to generate the third feature (S26).

[0087] The classification unit 26 performs identification prediction or classification prediction using the third feature in the classification block BL4 (S27). Specifically, the classification unit 26 inputs the third feature into the SoftMax classifier and performs identification prediction or classification prediction.

[0088] The data examples, numerical examples, model examples, number of data points, number of slices, number of models, and specific examples used in the above embodiment are merely examples and can be changed as desired.

[0089] The processing procedures, control procedures, specific names, various data, and parameters shown in the above documents and drawings may be changed at will unless otherwise specified.

[0090] Furthermore, the components of each illustrated device are functionally conceptual and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown. That is, all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions.

[0091] Furthermore, each processing function performed by each device can be implemented, in whole or in part, by a CPU and a program executed for analysis by that CPU, or by hardware using wired logic.

[0092] Figure 9 illustrates an example of a hardware configuration. As shown in Figure 9, the information processing device 10 includes a communication device 10a, an HDD (Hard Disk Drive) 10b, memory 10c, and a processor 10d. Furthermore, the components shown in Figure 9 are interconnected by a bus or the like.

[0093] The communication device 10a is a network interface card or the like, and communicates with other devices. The HDD 10b stores the program and database that operate the functions shown in Figure 5.

[0094] The processor 10d operates a process that performs the functions described in Figure 5 by reading a program that performs the same processing as each processing unit shown in Figure 5 from the HDD 10b or the like and loading it into memory 10c. For example, this process performs the same functions as each processing unit of the information processing device 10. Specifically, the processor 10d reads a program that has the same functions as the acquisition unit 21, preprocessing unit 22, first generation unit 23, extraction unit 24, second generation unit 25, classification unit 26, loss unit 27, and training unit 28 from the HDD 10b or the like. Then, the processor 10d executes a process that performs the same processing as the acquisition unit 21, preprocessing unit 22, first generation unit 23, extraction unit 24, second generation unit 25, classification unit 26, loss unit 27, and training unit 28.

[0095] Thus, the information processing device 10 operates as an information processing device that executes a machine learning method by reading and executing a program. Furthermore, the information processing device 10 can also achieve the same functionality as in the above-described embodiment by reading the program from a recording medium using a media reader and executing the read program. It should be noted that the program referred to in this other embodiment is not limited to being executed by the information processing device 10. For example, the present invention can be similarly applied when another computer or server executes the program, or when they collaborate to execute the program.

[0096] This program can be distributed via networks such as the internet. Furthermore, this program can be recorded on computer-readable storage media such as hard disks, flexible disks (floppy disks), CD-ROMs, MOs (Magneto-Optical disks), and DVDs (Digital Versatile Discs), and executed by reading the program from these media using a computer.

Claims

1. A machine learning method in which a computer performs a process to generate third features for input into a classifier that performs gait recognition, based on initial features generated from a first video of normal walking and first features extracted from a second video that includes both normal walking and walking with external attributes, from a dataset containing point cloud videos of gait recognition targets.

2. The machine learning method according to claim 1, wherein the generation process includes a process of generating the third feature based on the initial feature generated based on the first point cloud frame generated by performing preprocessing on the first video, and the first feature extracted from the second point cloud frame generated by performing the same preprocessing on the second video.

3. The machine learning method according to claim 1, wherein the generation process includes a process of generating a third feature based on initial features generated based on a predetermined number of first point cloud frames generated by performing preprocessing on each point cloud frame of the first video, and first features extracted from a predetermined number of second point cloud frames generated by performing the same preprocessing on each point cloud frame of the second video.

4. The machine learning method according to claim 1, wherein the generation process includes a process of generating a third feature based on an initial feature generated by aggregating a plurality of second features extracted from each video of the first video using a feature extractor, and the first feature extracted from the second video using the feature extractor.

5. The machine learning method according to claim 1, wherein the generation process includes a process of generating a third feature by aggregating the features that have been emphasized by mutual attention between the initial feature and the first feature, such that the temporal dynamics are captured.

6. The machine learning method according to claim 1, wherein the generation process includes a process of generating the third feature by aggregating the initial feature, which is generated by aggregating a plurality of second features extracted from each first point cloud frame of the first video using a feature extractor, and the features of each second point cloud frame, which are emphasized by mutual attention between the initial feature and a plurality of first features extracted from each second point cloud frame of the second video using the feature extractor, in such a way that the temporal dynamics are captured.

7. The machine learning method according to claim 1, wherein the computer further performs a process to train the feature extractor until the loss based on the gait recognition result using the third feature is minimized, the feature extractor is a feature extractor that outputs the third feature when the point cloud video is input, and is a feature extractor that is trained based on the error between the gait recognition result and the correct data of the gait recognition target.

8. The machine learning method according to claim 7, wherein the training process includes a process of training a feature extractor to generate the initial features based on the first video and to extract the first features from the second video.

9. The machine learning method according to claim 7, wherein the training process includes updating the initial features each time the number of training iterations for training the feature extractor reaches a predetermined threshold, and training the feature extractor using the latest updated initial features.

10. The machine learning method according to claim 1, which includes processing to determine that the object to be recognized as walking is a pedestrian.

11. The machine learning method according to claim 1, which includes processing that the external attribute is an item possessed by a pedestrian.

12. A computer-readable recording medium that records a machine learning program that causes a computer to perform a process to generate third features for input into a classifier that performs gait recognition, based on initial features generated from a first video of normal walking and first features extracted from a second video that includes both normal walking and walking with external attributes, from a dataset containing point cloud videos of gait recognition targets.

13. An information processing device including a control unit that performs a process to generate third features for input to a classifier that performs gait recognition, based on initial features generated from a first video of normal walking and first features extracted from a second video that includes both normal walking and walking with external attributes, from a dataset containing point cloud videos of gait recognition targets.