Label changing program, label changing method, and information processing device.

The information processing device enhances behavioral analysis accuracy by using machine learning models to correct target area settings based on individual movements and orientations, addressing inaccuracies in existing methods.

JP7864997B2Active Publication Date: 2026-05-26FUJITSU LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
FUJITSU LTD
Filing Date
2021-11-30
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing methods for setting target areas in behavioral analysis from video data, such as manual setting and semantic segmentation, suffer from inaccuracies leading to false positives and reduced accuracy in detecting purchasing behaviors.

Method used

An information processing device that uses a combination of machine learning models for semantic segmentation and motion analysis to identify and correct target areas by analyzing the orientation and movement of individuals within the video data, adjusting labels based on detected actions.

Benefits of technology

The solution effectively suppresses the deterioration of behavioral analysis accuracy by accurately setting target areas, reducing false positives and improving the precision of detecting relevant actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007864997000001
    Figure 0007864997000001
  • Figure 0007864997000002
    Figure 0007864997000002
  • Figure 0007864997000003
    Figure 0007864997000003
Patent Text Reader

Abstract

To suppress deterioration in the accuracy of behavior analysis.SOLUTION: An information processing apparatus acquires image data having a plurality of areas. The information processing apparatus sets a label for each of the plurality of areas by inputting the image data into a first machine learning model. The information processing apparatus identifies a behavior performed by a person located in a first area among the plurality of areas for an object located in a second area. The information processing apparatus changes a label set for the second area on the basis of the identified behavior of the person.SELECTED DRAWING: Figure 16
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a label change program, a label change method, and an information processing apparatus.

Background Art

[0002] Technological development for analyzing human behavior from video data captured by a camera has been underway. For example, from each image data included in the video data, a target area, which is an area where purchasing behavior is likely to occur, is extracted, and by detecting an operation of raising an arm to a certain position in the target area as a picking operation, the purchasing behavior is analyzed. In recent years, as a method for detecting the target area, for each image data, setting of the target area by manual operation or setting of the target area using semantic segmentation has been used.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the above technology, it is difficult to accurately set the target area. For example, in the manual method, it is necessary to set the target area for a huge amount of image data, which not only takes time but is also difficult to prevent human errors. Further, in the method using semantic segmentation, the entire passage where consumers walk in the store is set as the target area. For this reason, unnecessary picking operations are detected, and the accuracy of behavior analysis deteriorates.

[0005] One aspect of this project is to provide a label changing program, a label changing method, and an information processing device that can suppress the degradation of accuracy in behavioral analysis. [Means for solving the problem]

[0006] In the first proposal, the label changing program is characterized by causing a computer to acquire image data having multiple areas, input the image data into a first machine learning model to set a label for each of the multiple areas, identify an action taken by a person located in the first area of ​​the multiple areas toward an object located in the second area, and change the label set for the second area based on the identified action of the person. [Effects of the Invention]

[0007] According to one embodiment, it is possible to suppress the deterioration of the accuracy of behavioral analysis. [Brief explanation of the drawing]

[0008] [Figure 1] Figure 1 is a diagram illustrating the overall configuration of a system including an information processing device according to Example 1. [Figure 2] Figure 2 is a diagram illustrating the behavior of the recognized object in Example 1. [Figure 3] Figure 3 illustrates the detection of a region of interest using semantic segmentation. [Figure 4] Figure 4 illustrates the detection of the region of interest in the reference technology. [Figure 5] Figure 5 is a functional block diagram showing the functional configuration of the information processing device according to Embodiment 1. [Figure 6] Figure 6 illustrates the generation of the first machine learning model. [Figure 7] Figure 7 is a diagram illustrating the extraction processing unit according to Example 1. [Figure 8] Figure 8 illustrates the motion analysis using the second machine learning model. [Figure 9] FIG. 9 is a diagram for explaining the setting of a reference direction by tracking. [Figure 10] FIG. 10 is a diagram for explaining clustering. [Figure 11] FIG. 11 is a diagram for explaining the extraction of clusters. [Figure 12] FIG. 12 is a diagram for explaining the extraction of a region of interest. [Figure 13] FIG. 13 is a diagram for explaining the correction processing unit according to Example 1. [Figure 14] FIG. 14 is a diagram for explaining the setting of a reference line to the execution result of semantic segmentation. [Figure 15] FIG. 15 is a diagram for explaining clustering based on a reference line. [Figure 16] FIG. 16 is a diagram for explaining label correction. [Figure 17] FIG. 17 is a diagram for explaining the setting of a merchandise shelf area. [Figure 18] FIG. 18 is a flowchart showing the flow of the extraction process according to Example 1. [Figure 19] FIG. 19 is a flowchart showing the flow of the correction process according to Example 1. [Figure 20] FIG. 20 is a diagram for explaining the generation of a user's movement trajectory according to Example 2. [Figure 21] FIG. 21 is a diagram for explaining the plotting of a face direction and a body direction according to Example 2. [Figure 22] FIG. 22 is a diagram for explaining the extraction of a region of interest according to Example 2. [Figure 23] FIG. 23 is a diagram for explaining the generation of coordinates of a region of interest. [Figure 24] FIG. 24 is a flowchart showing the flow of the extraction process according to Example 2. [Figure 25] FIG. 25 is a diagram for explaining a hardware configuration example.

MODE FOR CARRYING OUT THE INVENTION

[0009] The following describes in detail, with reference to the drawings, embodiments of the label changing program, label changing method, and information processing device disclosed in this application. However, the present invention is not limited by these embodiments. Furthermore, each embodiment can be combined as appropriate within a non-consistent range. [Examples]

[0010] [Overall structure] Figure 1 is a diagram illustrating the overall configuration of a system including an information processing device 10 according to Embodiment 1. As shown in Figure 1, this system includes a store 1, which is an example of a space, a plurality of cameras 2 installed in different locations within the store 1, and an information processing device 10.

[0011] Each of the multiple cameras 2 is an example of a surveillance camera that captures a predetermined area within the store 1, and transmits the captured video data to the information processing device 100. In the following explanation, the video data may be referred to as "video data." The video data also contains multiple image frames in chronological order. Each image frame is assigned a frame number in ascending order of chronological order. A single image frame is image data of a still image captured by camera 2 at a certain point in time.

[0012] The information processing device 10 is an example of a computer that analyzes the image data captured by each of the multiple cameras 2. Each of the multiple cameras 2 and the information processing device 10 are connected using various networks, such as the internet or dedicated lines, whether wired or wireless.

[0013] In recent years, technological development has progressed in analyzing human behavior from video data captured by camera 2. For example, by extracting attention regions—areas where purchasing behavior is likely to occur—from each image data contained in the video data, and detecting the action of raising the arm to a certain position within the attention region as a picking action, purchasing behavior can be analyzed.

[0014] Figure 2 illustrates the actions of the object to be recognized in Example 1. Let's assume that region A in Figure 2 is the area of ​​interest. In this case, the object to be recognized is the picking action of a person (user) located in front of a product shelf, as shown in Figure 2(1). However, false positives occur because a person performing an action similar to picking, but without reaching for a product, in a location without a product shelf, as shown in Figure 2(2) and (3), is also recognized.

[0015] To reduce false positives, one might consider setting the area of ​​the product shelf that a person is reaching for as the area of ​​focus. However, in that case, the person shown in (2) of Figure 2 would be falsely detected because their hand is on the shelf in the image data. Another method is to set the aisle at the person's feet as the area of ​​focus. For example, if area A shown in Figure 1 is set as the area of ​​focus, the people shown in (2) and (3) of Figure 2 will not be detected, and false positives will be suppressed.

[0016] The method of setting the area at the user's feet as the area of ​​focus is often done manually. However, manual setting involves setting the area of ​​focus for a vast amount of image data, which is not only time-consuming but also makes it difficult to prevent human error.

[0017] Another method used is automatic assignment through semantic segmentation, a technique that categorizes what is depicted in each pixel of image data. Figure 3 illustrates the detection of areas of interest using semantic segmentation. As shown in Figure 3, semantic segmentation inputs image data into a machine learning model (convolutional encoder-decoder) and obtains an output result in which each region of the image data is labeled. However, the label "pathway" is assigned not only to the areas of interest but also to all pathways, including those outside the areas of interest. As a result, the people in (2) and (3) in Figure 2 are also included in the recognition target, leading to false detections of picking actions by these people.

[0018] Furthermore, it is also possible to use a reference technology that extracts people's work positions from camera video data and automatically provides ROI (Region of Interest) through clustering of work positions. Figure 4 is a diagram illustrating the detection of the area of ​​interest in the reference technology. As shown in Figure 4, the reference technology extracts areas where stationary purchasing behavior occurs, so only the stationary positions of the person shown in Figure 4(B) are extracted, making it difficult to adequately cover the area of ​​interest A. In other words, the reference technology has difficulty detecting the action of slowly moving to pick up an item (picking action).

[0019] Generally, within image data, areas designated as areas of interest or areas where picking actions are detected are defined as product shelf areas where products picked up by individuals are displayed, and then human behavior analysis is performed. However, as mentioned above, it is difficult to accurately extract areas of interest using manual settings, semantic segmentation, or reference techniques. This can lead to errors in defining product shelf areas, resulting in a deterioration of the accuracy of the final behavior analysis.

[0020] Therefore, the information processing device 10 according to Embodiment 1 acquires image data having multiple areas, inputs the image data into a machine learning model to perform Symantec segmentation, and sets a label for each of the multiple areas. The information processing device 10 identifies an action taken by a person located in one of the first areas of the multiple areas towards an object located in a second area, and changes the label set for the second area based on the action of the identified person.

[0021] In other words, the information processing device 10 uses the fact that purchasing behavior in retail stores mainly involves movement and product selection, and that there is variation in the orientation of the body relative to the aisle direction when selecting products, to extract areas of interest where product selection occurs and to correct the segmentation results. As a result, the information processing device 10 can accurately set areas of interest that are the target of behavioral analysis.

[0022] [Functional Configuration] Figure 5 is a functional block diagram showing the functional configuration of the information processing device 10 according to Embodiment 1. As shown in Figure 5, the information processing device 10 has a communication unit 11, a storage unit 12, and a control unit 20.

[0023] The communication unit 11 is a processing unit that controls communication with other devices, and is implemented, for example, by a communication interface. For example, the communication unit 11 receives video data from the camera 2 and transmits the processing results from the control unit 20 to a management terminal or the like.

[0024] The memory unit 12 is a processing unit that stores various data and programs executed by the control unit 20, and is implemented using memory or a hard disk. The memory unit 12 stores training data DB 13, first machine learning model 14, second machine learning model 15, video data DB 16, segment result DB 17, ROI information DB 18, and setting result DB 19.

[0025] The training data DB13 is a database that stores each training data used to train the first machine learning model 14. Specifically, each training data is data that associates RGB image data, which is the explanatory variable, with the result of performing semantic segmentation on that image data (hereinafter sometimes referred to as segment result or segmentation result), which is the target variable (ground truth information).

[0026] The first machine learning model 14 is a model that performs semantic segmentation. Specifically, the first machine learning model 14 outputs segmentation results in response to RGB image data input. The segmentation results are assigned identified labels to each region within the image data. For example, the first machine learning model 14 can employ a convolutional encoder-decoder.

[0027] The second machine learning model 15 is a model that performs motion analysis. Specifically, the second machine learning model 15 is a pre-trained model that estimates the 2D joint positions (skeletal coordinates) of the head, wrists, hips, ankles, etc., from 2D image data of a person, and is an example of a deep learning machine that recognizes basic movements and user-defined rules. By using this second machine learning model 15, it is possible to recognize the basic movements of a person and obtain the position of the ankles, the direction of the face, and the orientation of the body.

[0028] The video data DB16 is a database that stores video data captured by each of the multiple cameras 2 installed in store 1. For example, the video data DB16 stores video data for each camera 2, or for each time period in which the video was captured.

[0029] The segment results DB17 is a database that stores the results of semantic segmentation. Specifically, the segment results DB17 stores the output results of the first machine learning model 14. For example, the segment results DB17 stores RGB image data in association with the results of semantic segmentation.

[0030] The ROI information DB18 is a database that stores ROIs of areas of interest, ROIs of product shelves, etc., obtained by the control unit 20, which will be described later. For example, the ROI information DB18 stores the ROIs of areas of interest, ROIs of product shelves, etc., associated with each RGB image data.

[0031] The setting result DB19 is a database that stores the results of setting product shelf areas for segment results by the control unit 20, which will be described later. For example, the setting result DB19 stores RGB image data and the setting information for each label set for the image data in association with each other.

[0032] The control unit 20 is the processing unit that oversees the entire information processing device 10, and is implemented by, for example, a processor. This control unit 20 includes a pre-learning unit 30, an acquisition unit 40, an extraction processing unit 50, a correction processing unit 60, and an area setting unit 70. The pre-learning unit 30, acquisition unit 40, extraction processing unit 50, correction processing unit 60, and area setting unit 70 are implemented by electronic circuits of the processor or processes executed by the processor.

[0033] The pre-training unit 30 is a processing unit that generates the first machine learning model 14. Specifically, the pre-training unit 30 trains the first machine learning model 14 using machine learning with each training data stored in the training data DB 13.

[0034] Figure 6 illustrates the generation of the first machine learning model 14. As shown in Figure 6, the pre-training unit 30 inputs training data, including RGB image data and ground truth information (segmentation results), into the first machine learning model 14 and obtains output results (segmentation results). The pre-training unit 30 then optimizes the parameters of the first machine learning model 14 so as to minimize the error between the ground truth information of the training data and the output results.

[0035] The acquisition unit 40 is a processing unit that acquires video data from each camera 2 and stores it in the video data DB 16. For example, the acquisition unit 40 may acquire data from each camera 2 as needed, or it may acquire data periodically.

[0036] The extraction processing unit 50 is a processing unit that extracts areas of interest within video data by performing motion analysis on the video data. Figure 7 is a diagram illustrating the extraction processing unit 50 according to Embodiment 1. As shown in Figure 7, the extraction processing unit 50 includes a tracking unit 51, a motion analysis unit 52, a reference line extraction unit 53, a clustering execution unit 54, an angle calculation unit 55, and an area of ​​interest extraction unit 56.

[0037] The tracking unit 51 is a processing unit that performs tracking of the same person in the video data. For example, the tracking unit 51 uses known tracking technology to track the same person in the video data and extracts the person's movement path (movement trajectory). The tracking unit 51 then outputs the extraction result to the reference line extraction unit 53, etc.

[0038] The motion analysis unit 52 is a processing unit that performs motion analysis of people captured in video data by the camera 2. Specifically, the motion analysis unit 52 inputs each image data (frame) contained in the video data into the second machine learning model 15 and recognizes the motion of the person captured in each image data.

[0039] Figure 8 illustrates motion analysis using the second machine learning model 15. As shown in Figure 8, the motion analysis unit 52 inputs RGB image data into the second machine learning model 15 and obtains the two-dimensional skeletal coordinates of the person in the image data. The motion analysis unit 52 then identifies the position of the person's ankles, the orientation of their face, and the orientation of their body according to the two-dimensional skeletal coordinates, and outputs the identified results to the clustering execution unit 54 or the like.

[0040] In this way, the motion analysis unit 52 inputs each image data (for example, 100 frames) contained in each video data acquired at predetermined time intervals into the second machine learning model 15, and by measuring the position of the ankles, the orientation of the face, and the orientation of the body of the person shown in each image data, it can identify the transitions in the position of the ankles, the orientation of the face, and the orientation of the body of the person within the video data.

[0041] The reference line extraction unit 53 is a processing unit that extracts a person's movement path from tracking information and sets the direction of the path to be used as a reference line. Specifically, the reference line extraction unit 24 acquires (selects) image data from the video data and uses the movement path of a person obtained by the tracking unit 51 to set a reference direction, which is the direction in which the user is walking, on the acquired image data. The reference line extraction unit 53 then extracts the set reference direction as a reference line indicating the movement path. The reference line extraction unit 53 can select any image data from the video data, such as the first or last image data of the video data, as the image data.

[0042] Figure 9 illustrates the setting of a reference direction by tracking. As shown in Figure 9, the reference line extraction unit 53 sets the tracking results, movement paths A1 and A2, on the image data. At this time, the reference line extraction unit 53 can set the area containing the set movement paths as the area of ​​the pathway. The reference line extraction unit 53 can set the area of ​​the pathway on the image data based on the results of performing semantic segmentation on the image data.

[0043] Next, the reference line extraction unit 53 identifies the transition from movement path A1 to movement path A2 based on the tracking results, and sets reference directions B1, B2, and B3 on the area of ​​the passage according to that transition. The reference line extraction unit 53 then sets each of these reference directions B1, B2, and B3 as reference lines. Note that movement paths and transitions between movement paths are not limited to one direction, but may be identified in multiple directions. Even in this case, if the movement trajectory is the same excluding the direction, it is considered one passage direction and is extracted as one reference line. For example, the reference line extraction unit 53 calculates an approximate straight line that represents the passage direction from multiple movement paths walked by the user, and sets that approximate curve as a reference line. The reference line extraction unit 53 also outputs the result of setting the reference lines to the clustering execution unit 54, etc.

[0044] The clustering execution unit 54 is a processing unit that extracts the movement trajectory of each person and generates multiple clusters by clustering based on the distance between each reference line and each person's movement trajectory. Specifically, the clustering execution unit 54 clusters each movement trajectory based on which reference line it is closest to.

[0045] Figure 10 is a diagram illustrating clustering. As shown in Figure 10, the clustering execution unit 54 obtains the position of the ankle of each person in the image data from the motion analysis unit 52 and plots it on the image data where reference lines B1, B2, and B3 have been set. The clustering execution unit 54 then generates multiple clusters by clustering based on the distance between each reference line and the movement trajectory of each person.

[0046] For example, the clustering execution unit 54 draws perpendiculars from each movement trajectory to each reference line and performs clustering based on the length of these perpendiculars, thereby clustering each movement trajectory to one of the reference lines. Note that the base distance is not limited to the length of the perpendiculars; Euclidean distance or other distances can also be used.

[0047] As a result, the clustering execution unit 54 generates cluster C1 containing the point cloud of the movement trajectory closest to the reference line B1, cluster C2 containing the point cloud of the movement trajectory closest to the reference line B2, and cluster C3 containing the point cloud of the movement trajectory closest to the reference line B3. The clustering execution unit 54 then outputs the clustering results to the angle calculation unit 55 and the like.

[0048] The angle calculation unit 55 is a processing unit that calculates the angle between the body orientation and each reference line for each clustered movement trajectory. Specifically, the angle calculation unit 55 obtains the body orientation of the person in each image data from the motion analysis unit 52 and associates the corresponding body orientation with the movement trajectory in the image data. Then, the angle calculation unit 55 uses the clustering results to identify the reference line of the cluster to which each movement trajectory belongs. After that, the angle calculation unit 55 calculates the angle between the body orientation and the reference line of the cluster to which each movement trajectory belongs using a known method. Note that the angle calculation unit 55 can also use the orientation of the face, not just the body orientation. The angle calculation unit 55 outputs the angle corresponding to each movement trajectory to the area of ​​interest extraction unit 56, etc.

[0049] The focus region extraction unit 56 is a processing unit that, for each of the multiple clusters, extracts regions containing clusters where the evaluation value based on the angle between each movement trajectory belonging to the cluster and the reference line is greater than or equal to a threshold. Specifically, the focus region extraction unit 56 extracts reference lines that contain many large angles among the angles that the body orientation makes with each reference line, and extracts the regions to which such reference lines belong as focus regions.

[0050] Figure 11 illustrates the extraction of clusters. As shown in Figure 11, the area of ​​interest extraction unit 56 plots the body orientation corresponding to each movement trajectory for each movement trajectory plotted in the image data. The area of ​​interest extraction unit 56 also associates the calculated angle with each movement trajectory.

[0051] The area of ​​interest extraction unit 56 then aggregates the angles of the movement trajectories belonging to each cluster. For example, as shown in Figure 11, the area of ​​interest extraction unit 56 aggregates the angles of each movement trajectory belonging to cluster C1 and the number of movement trajectories corresponding to those angles, the angles of each movement trajectory belonging to cluster C2 and the number of movement trajectories corresponding to those angles, and the angles of each movement trajectory belonging to cluster C3 and the number of movement trajectories corresponding to those angles.

[0052] Subsequently, the area of ​​interest extraction unit 56 extracts clusters that have many large angles. For example, the area of ​​interest extraction unit 56 calculates evaluation values ​​for each cluster, such as the median angle, the average angle, and the percentage of angles greater than or equal to 60 degrees. Then, the area of ​​interest extraction unit 56 extracts clusters C2 and C3, whose evaluation values ​​are above a threshold.

[0053] Next, the focus region extraction unit 56 generates polygons that enclose each movement trajectory belonging to the extracted clusters C2 and C3 as focus regions. Figure 12 is a diagram illustrating the extraction of focus regions. As shown in Figure 12, for cluster C2, the focus region extraction unit 56 generates the largest polygon C2' that includes each movement trajectory belonging to cluster C2 and extracts it as a focus region. Similarly, for cluster C3, the focus region extraction unit 56 generates the largest polygon C3' that includes each movement trajectory belonging to cluster C3 and extracts it as a focus region.

[0054] Furthermore, the area of ​​interest extraction unit 56 stores the coordinates of each polygon in the ROI information DB 18 or outputs them to the area setting unit 28. The area of ​​interest extraction unit 56 also stores information related to the set area of ​​interest, such as image data in which the area of ​​interest has been set, in the setting result DB 19.

[0055] Returning to Figure 5, the correction processing unit 60 is a processing unit that uses the extraction results of the extraction processing unit 50 to correct (change) the labels of each area obtained by semantic segmentation. Figure 13 is a diagram illustrating the correction processing unit 60 according to Embodiment 1. As shown in Figure 13, the correction processing unit 60 includes an extraction result acquisition unit 61, a semantic segmentation unit 62, a baseline setting unit 63, a clustering execution unit 64, and a label correction unit 65.

[0056] The extraction result acquisition unit 61 is a processing unit that acquires the processing results of the extraction processing unit 50. For example, the extraction result acquisition unit 61 acquires information about the baseline, extraction results of the region of interest, information about ROI, and behavior recognition results such as the position of the ankle, the orientation of the body, and the orientation of the face from the correction processing unit 60, and outputs them to the baseline setting unit 63, the clustering execution unit 64, etc.

[0057] The semantic segmentation unit 62 is a processing unit that assigns labels to each area of ​​image data using semantic segmentation. For example, the semantic segmentation unit 62 inputs image data included in the video data, such as the image data used for extracting the region of interest by the extraction processing unit 50, into the first machine learning model 14. The semantic segmentation unit 62 then obtains the results of the semantic segmentation performed by the first machine learning model 14.

[0058] The semantic segmentation unit 62 outputs the semantic segmentation result (segmentation result) to the reference line setting unit 63. The segmentation result includes labels indicating the identified result for each of the multiple regions included in the image data. For example, the semantic segmentation result may include labels such as "shelf," "aisle," and "wall."

[0059] The reference line setting unit 63 is a processing unit that sets reference lines in the segmentation results. Figure 14 is a diagram illustrating the setting of reference lines in the results of semantic segmentation. As shown in Figure 14, the reference line setting unit 63 obtains the segmentation results from the semantic segmentation unit 62 and obtains information about the reference lines from the extraction result acquisition unit 61. Then, the reference line setting unit 63 plots reference lines B1, B2, and B3 on the segmentation results.

[0060] The clustering execution unit 64 is a processing unit that performs clustering based on the reference lines set by the reference line setting unit 63 on the segmentation results. Figure 15 is a diagram illustrating clustering based on reference lines. As shown in Figure 15, the clustering execution unit 64 identifies the area labeled "pathway" among the labels set (identified) in the segmentation results. The clustering execution unit 64 then calculates the distance between each pixel belonging to the identified "pathway" area and each reference line (B1, B2, B3), and clusters each pixel so that it belongs to the reference line with the closest distance. The distance can be calculated using the length of the perpendicular from each pixel to each reference line, or the Euclidean distance between the pixel and the reference line.

[0061] The clustering execution unit 64 then identifies cluster L1 belonging to reference line B1, cluster L2 belonging to reference line B2, and cluster L3 belonging to reference line B3. Subsequently, the clustering execution unit 64 outputs the identification results, etc., to the label correction unit 65, etc.

[0062] The label modification unit 65 is a processing unit that modifies the labels of the segmentation results based on the extraction results of the extraction processing unit 50. Specifically, the label modification unit 65 identifies a cluster of interest that corresponds to the area of ​​interest among multiple clusters, modifies the area of ​​the cluster of interest to include the area of ​​interest, and changes the label set for the modified area to a label corresponding to the area of ​​interest. In other words, the label modification unit 65 modifies the area of ​​each cluster so that the area containing the clustering results generated by the clustering execution unit 64 and the area of ​​interest extracted by the extraction processing unit 50 is maximized, and labels the modified area as the area of ​​interest.

[0063] Figure 16 illustrates the label correction process. As shown in Figure 16, the label correction unit 65 obtains the coordinates of each polygon related to the region of interest (C2' and C3') from the extraction processing unit 50 and maps them to the clustered segmentation result (image data). The label correction unit 65 then identifies cluster L2 to which the region of interest C2' belongs and cluster L3 to which the region of interest C3' belongs.

[0064] Subsequently, the label modification unit 65 generates a region L2' by expanding the region of cluster L2 so that it includes the region of interest C2'. Then, the label modification unit 65 modifies (changes) the label "pathway" set in region L2' to the label "region of interest".

[0065] Similarly, the label modification unit 65 generates a region L3' by extending the region of cluster L3 so that it includes the region of interest C3'. Then, the label modification unit 65 modifies the label "pathway" set in region L3' to the label "region of interest".

[0066] Furthermore, if the area of ​​interest is larger than the cluster area, the label correction unit 65 corrects (changes) the label "pathway" for the area of ​​interest to the label "area of ​​interest". The label correction unit 65 outputs the segmentation result with the corrected labels to the area setting unit 70.

[0067] Returning to Figure 5, the area setting unit 70 is a processing unit that sets, based on the orientation of the face or body, the area among the multiple areas that make up store 1 that is adjacent to the label "area of ​​interest" and where objects related to the person are stored. Specifically, the area setting unit 28 identifies the product shelf area where the product to be picked is located from the image data. That is, the area setting unit 70 changes the labels that have already been set based on the segmentation results for areas adjacent to area L2' and area L3' to the label "product shelf".

[0068] Figure 17 is a diagram illustrating the setting of product shelf areas. As shown in Figure 17, the area setting unit 70 obtains and plots, from the extraction processing unit 50, each of the movement trajectories and face orientations belonging to each area L2' and L3', which are labeled as "areas of interest".

[0069] The area setting unit 70 then identifies a direction in which the number of face orientation vectors is greater than or equal to a threshold, and identifies regions E1 and E2 as regions that are adjacent to or touch region L2' among the regions in that direction. As a result, the area setting unit 28 sets the labels of regions E1 and E2 to "product shelf area" in the segmentation results.

[0070] Similarly, the area setting unit 70 identifies a direction in which the number of face orientation vectors is greater than or equal to a threshold, and identifies regions E3 and E4 as regions that are adjacent to or touch region L3' among the regions in that direction. As a result, the area setting unit 28 sets the labels of regions E3 and E4 as "product shelf area" in the segmentation results.

[0071] The area setting unit 70 then stores information such as the coordinates of areas E1, E2, E3, and E4, as well as image data for each of areas E1 through E4, in the setting result DB 19. The area setting unit 70 can also set areas for the "product shelf area" based on the image data that formed the basis of the segmentation results, rather than the segmentation results themselves.

[0072] [Extraction Process Flow] Figure 18 is a flowchart showing the flow of the extraction process according to Example 1. As shown in Figure 18, when the extraction processing unit 50 is instructed to start processing (S101: Yes), the extraction processing unit 50 acquires video data from the video data DB 16 (S102).

[0073] Next, the extraction processing unit 50 performs person tracking based on the video data (S103), and sets a reference direction based on the person tracking results (S104). For example, the extraction processing unit 50 tracks the same person in the video data to extract their movement path and sets a reference line using the movement path the user walks.

[0074] Furthermore, the extraction processing unit 50 performs behavioral analysis using each image data that makes up the video data (S105), and obtains the position and orientation of the person based on the results of the behavioral analysis (S106). For example, the extraction processing unit 50 uses the second machine learning model 15 to identify the orientation of the face, the orientation of the body, the position of the ankles, and the transitions between these for each person in the video data.

[0075] Subsequently, the extraction processing unit 50 extracts the movement trajectory of each person and generates multiple clusters by clustering based on the distance between each reference line and each person's movement trajectory (S107). For example, the extraction processing unit 50 clusters each movement trajectory based on which reference line it is closest to.

[0076] Next, the extraction processing unit 50 calculates an angle for each cluster (S108). For example, the extraction processing unit 50 calculates the angle between the orientation of the body corresponding to each movement trajectory and the reference line of the cluster to which each movement trajectory belongs.

[0077] Then, the extraction processing unit 50 calculates the median angle of each movement trajectory belonging to each cluster (S109) and extracts clusters whose median is greater than or equal to a threshold (S110). Subsequently, the extraction processing unit 50 generates a polygonal region that surrounds (includes) all movement trajectories belonging to the extracted clusters and extracts this region as the region of interest (S111).

[0078] Subsequently, the extraction processing unit 50 outputs the information obtained from the extraction process, such as information on the area of ​​interest, the coordinates of the polygon, and the action recognition results, to the storage unit 12 and the modification processing unit 60 (S112).

[0079] [Correction Process Flow] Figure 19 is a flowchart showing the flow of the correction process according to Example 1. As shown in Figure 19, the correction processing unit 60 acquires the information obtained in the extraction process in Figure 18 (S201), inputs the image data into the first machine learning model 14, and obtains the result of performing semantic segmentation of the image data (S202).

[0080] Next, the correction processing unit 60 plots a baseline on the result of the semantic segmentation (S203) and performs clustering based on the baseline (S204). For example, the correction processing unit 60 clusters each pixel in the pathway to determine which baseline it is closest to.

[0081] Then, the correction processing unit 60 superimposes the extraction results onto the clustering results (S205). For example, the correction processing unit 60 maps the polygon of the region of interest generated in the process shown in Figure 18 onto the clustering results.

[0082] Subsequently, the correction processing unit 60 performs label correction based on the superimposition result (S206). For example, the correction processing unit 60 expands the cluster area to include the area of ​​interest to the maximum extent, and corrects the label "aisle" of the area to which the expanded area belongs to the label "area of ​​interest". Then, the area setting unit 28 sets the product shelf area adjacent to the area of ​​interest based on the orientation of the face or body (S207).

[0083] [effect] As described above, the information processing device 10 performs semantic segmentation to divide the image data into regions, re-extracts the aisle regions from the segmentation results and motion analysis results, extracts the variability of face orientation and body orientation from the motion analysis results, and extracts the region of interest from the aisle regions and variability information by clustering. The information processing device 10 then performs clustering on the aisle regions of the segmentation results, modifies the region to take the maximum value between the clustering results and the extracted region of interest, and labels the modified region as the region of interest.

[0084] As a result, the information processing device 10 can suppress the problem of excess or deficiency in the extracted regions when attempting to extract regions of interest, and can automatically provide regions of interest without excess or deficiency. Therefore, the information processing device 10 can accurately set the regions of interest to be targeted for behavioral analysis. [Examples]

[0085] By the way, in Example 1, we described an example in which a baseline is extracted and a region of interest (coordinates of a polygon) is generated by clustering using the baseline. However, the extraction of the region of interest is not limited to this. For example, the information processing device 10 can also extract the region of interest by using the fact that when moving, the face and body are facing the same direction, but when making a selection, there is variation in the orientation of the face and body.

[0086] Therefore, in Example 2, we will describe an example in which the extraction processing unit 50 performs a separate process to extract the region of interest using variations in the orientation of the face and body. Note that the processing of the correction processing unit 60 is the same as in Example 1, so a detailed explanation will be omitted.

[0087] First, the extraction processing unit 50 inputs each image data (frame) contained in the video data captured by the camera 2 into the second machine learning model 15 and recognizes the movements of the person in each image data. Specifically, the extraction processing unit 50 identifies the person's two-dimensional skeletal coordinates, the position of the person's ankles, the orientation of the face, the orientation of the body, etc., using the method described in Figure 8.

[0088] For example, the extraction processing unit 50 inputs each image data (e.g., 100 frames) contained in each video data acquired at predetermined time intervals into the second machine learning model 15, and by measuring the position of the ankles, the orientation of the face, and the orientation of the body of the person shown in each image data, it can identify the transitions in the position of the ankles, the orientation of the face, and the orientation of the body of the person within the video data.

[0089] Next, the extraction processing unit 50 uses the person's two-dimensional skeletal coordinates to extract the variation between the orientation of the person's body and the orientation of their face. Specifically, the extraction processing unit 50 obtains the orientation of the face and the orientation of the body for each image data (for example, 100 frames) contained in the video data from the motion analysis unit 23. Subsequently, the extraction processing unit 50 calculates the angle between the orientation of the person's face and the orientation of their face in each image data as variation.

[0090] Next, the extraction processing unit 50 generates the movement trajectory of each person in the video data. Specifically, the extraction processing unit 50 generates the movement trajectory of each person by plotting the position of the person's ankle on the result of performing semantic segmentation on the image data within the video data.

[0091] Figure 20 illustrates the generation of the user's movement trajectory according to Embodiment 2. As shown in Figure 20, the extraction processing unit 50 inputs image data (for example, the last image data) from the video data to the first machine learning model 14. The extraction processing unit 50 then obtains segmentation results in which regions (areas) are identified by the first machine learning model 14 and labels are set for each area.

[0092] Subsequently, the extraction processing unit 50 identifies the area of ​​the passage labeled "passage" from each label included in the segmentation result. Next, the extraction processing unit 50 plots the position of each person's ankle, identified from each image data within the video data, as a trajectory within the passage area. In this way, the extraction processing unit 50 can generate a movement trajectory for the video data, showing how the people appearing in the video data move through the passage area.

[0093] Next, the extraction processing unit 50 extracts regions from the generated movement trajectories that include movement trajectories in which the angle between the orientation of the person's face and the orientation of the face is greater than or equal to a threshold, as regions of interest. Figure 21 is a diagram illustrating the plot of face orientation and body orientation according to Example 2, and Figure 22 is a diagram illustrating the extraction of regions of interest according to Example 2.

[0094] As shown in Figure 21, the extraction processing unit 50 plots the orientation of the identified person's face and body on the generated movement trajectory. Next, based on the calculated angle (variability), the extraction processing unit 50 identifies the angles of the person's face and body orientation for each trajectory. Then, as shown in Figure 22, the extraction processing unit 50 performs clustering on the point cloud of movement trajectories based on the variability between the face and body orientation. The extraction processing unit 50 then extracts regions M1 and M2, which are clustered as having angles above a threshold and large variability, as regions of interest, and extracts region M3, which is clustered as having angles below a threshold and small variability, as a pathway region.

[0095] Finally, the extraction processing unit 50 generates the coordinates of the region of interest. Figure 23 illustrates the generation of coordinates of the region of interest. As shown in Figure 23, the extraction processing unit 50 generates a polygon G surrounding the trajectories (point cloud) belonging to cluster M1, which has been extracted as the region of interest, and extracts the coordinates of polygon G. Similarly, the extraction processing unit 50 generates a polygon H surrounding the trajectories belonging to cluster M2, which has been extracted as the region of interest, and extracts the coordinates of polygon H.

[0096] In this way, the extraction processing unit 50 can narrow down the areas of interest within the video data that are subject to human behavior analysis and are areas where picking actions on products are detected. The correction processing unit 60 uses the information of the areas of interest (e.g., polygon coordinates) generated by the method described in Example 2 to perform the processing shown in Figure 19. Note that the extraction processing unit 50 is not limited to clustering; for example, it can also use methods such as extracting areas of interest that contain the most trajectories with angles greater than or equal to a threshold.

[0097] [Process flow] Figure 24 is a flowchart showing the flow of the extraction process according to Example 2. As shown in Figure 24, when the start of processing is instructed (S301: Yes), the extraction processing unit 50 performs motion analysis based on the video data (S302). Then, the extraction processing unit 50 detects the orientation of the person's face, etc., based on the motion analysis (S303). For example, the extraction processing unit 50 inputs each image data contained in the video data into the second machine learning model 15 to identify the two-dimensional skeletal information of the person contained in each image data and the transition of the two-dimensional skeletal information, and detects the position of the person's ankle, the orientation of their face, and the orientation of their body.

[0098] Next, the extraction processing unit 50 inputs the image data contained in the video data into the first machine learning model 14 and obtains the segmentation result, which is the result of performing semantic segmentation (S304).

[0099] The extraction processing unit 50 then generates the movement trajectory of each person from each image data contained in the video data (S305). For example, the extraction processing unit 50 generates the movement trajectory of each person by plotting the position of the ankle identified for each person in each image data on the segmentation result.

[0100] Subsequently, the extraction processing unit 50 plots the face orientation and body orientation for each movement trajectory in the segmentation result where the movement trajectories are plotted (S306). Then, the extraction processing unit 50 detects the variation between the face orientation and body orientation (S307). For example, for each movement trajectory, the extraction processing unit 50 obtains the angle between the vectors of the face orientation and body orientation as the variation.

[0101] Next, the extraction processing unit 50 performs clustering based on the variation between the orientation of the face and the orientation of the body (S308), and extracts areas of interest based on the clustering results (S309). For example, the area of ​​interest extraction unit 26 extracts clusters of trajectories where the angle is greater than or equal to a threshold as areas of interest. After that, the extraction processing unit 50 outputs the information obtained from the extraction process, such as information on areas of interest, polygon coordinates, and action recognition results, to the storage unit 12 and the modification processing unit 60 (S310).

[0102] [effect] By using this information processing device 10, there is no need to manually set the area of ​​interest, thus reducing human error. Compared to manual setting, it is possible to set the area of ​​interest accurately and quickly for a large amount of image data. Furthermore, since the information processing device 10 can extract the area in which a person moves their face to indicate interest as the area of ​​interest, it is possible to set the area of ​​interest without excess or deficiency, unlike the reference technology in Figure 4.

[0103] Furthermore, since the information processing device 10 can identify the area of ​​interest and adjacent areas as product shelves without any excess or deficiency, unlike the reference technology, it can detect picking operations where the user moves slowly to pick up products, not just picking operations when the user is stationary. As a result, the information processing device 10 can improve the accuracy of picking operation detection and thus improve the accuracy of behavioral analysis. [Examples]

[0104] Now, although embodiments of the present invention have been described, the present invention may be implemented in various other forms besides those described above.

[0105] [Numerical values, etc.] The numerical examples, number of cameras, label names, and number of trajectories used in the above embodiment are merely examples and can be changed as needed. Furthermore, the processing flow described in each flowchart can also be modified as appropriate within a consistent range. While the above embodiment uses a store as an example, it is not limited to this and can be applied to warehouses, factories, classrooms, train cars, airplane cabins, etc. In these cases, instead of the product shelf area described as an example of an area where objects related to people are stored, areas where objects are placed or luggage is stored will be detected and set.

[0106] Furthermore, while the above embodiment describes an example using the position of a person's ankle, it is not limited to this, and other positions such as the foot or shoe can also be used. Also, while the above embodiment describes an example of identifying a product shelf area in the direction of the face's orientation, it is also possible to identify a product shelf area in the direction of the body's orientation. In addition, each machine learning model can utilize a neural network or the like.

[0107] [system] Unless otherwise specified, the processing procedures, control procedures, specific names, and various data and parameters shown in the above documents and drawings may be changed at will.

[0108] Furthermore, the components of each illustrated device are functionally conceptual and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown. That is, all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions.

[0109] Furthermore, each processing function performed by each device may be implemented, in whole or in part, by a CPU and a program executed for analysis by that CPU, or by hardware using wired logic.

[0110] [Hardware] Figure 25 illustrates an example of a hardware configuration. As shown in Figure 25, the information processing device 10 includes a communication device 10a, an HDD (Hard Disk Drive) 10b, memory 10c, and a processor 10d. Furthermore, the components shown in Figure 25 are interconnected by a bus or the like.

[0111] The communication device 10a is a network interface card or the like, and communicates with other devices. The HDD 10b stores programs and databases that operate the functions shown in Figure 5.

[0112] The processor 10d operates a process that performs the functions described in Figure 5 by reading a program that performs the same processing as each processing unit shown in Figure 5 from the HDD 10b or the like and loading it into memory 10c. For example, this process performs the same functions as each processing unit of the information processing device 10. Specifically, the processor 10d reads a program that has the same functions as the pre-learning unit 30, acquisition unit 40, extraction processing unit 50, correction processing unit 60, area setting unit 70, etc. from the HDD 10b or the like. Then, the processor 10d executes a process that performs the same processing as the pre-learning unit 30, acquisition unit 40, extraction processing unit 50, correction processing unit 60, area setting unit 70, etc.

[0113] Thus, the information processing device 10 operates as an information processing device that executes an information processing method by reading and executing a program. Furthermore, the information processing device 10 can also achieve the same functionality as the embodiment described above by reading the program from the recording medium using a media reader and executing the read program. Note that the program referred to in this other embodiment is not limited to being executed by the information processing device 10. For example, the above embodiment may also be applied to cases where another computer or server executes the program, or where these computers or servers collaborate to execute the program.

[0114] This program may be distributed via a network such as the Internet. Alternatively, this program may be recorded on a computer-readable recording medium such as a hard disk, flexible disk (FD), CD-ROM, MO (Magneto-Optical disk), or DVD (Digital Versatile Disc), and executed by being read from the recording medium by a computer. [Explanation of Symbols]

[0115] 10 Information Processing Devices 11 Communications Department 12 Storage section 13 Training Data Database 14. First Machine Learning Model 15. Second Machine Learning Model 16 Video Data Database 17 Segment Results DB 18 ROI information DB 19 Configuration Result DB 20 Control Unit 30 Pre-learning section 40 Acquisition Department 50 Extraction Processing Unit 51 Tracking Section 52 Motion analysis section 53 Reference line extraction part 54 Clustering Execution Unit 55 Angle Calculation Unit 56 Attention area extraction part 60 Correction Processing Unit 61 Extraction result acquisition part 62 Semantic Segmentation Section 63 Reference line setting section 64 Clustering Execution Unit 65 Label Correction Section 70 Area setting section

Claims

1. On the computer, Acquire image data with multiple areas, By inputting the aforementioned image data into the first machine learning model, a label is set for each of the multiple areas. Identify the actions taken by a person located in the first of the aforementioned multiple areas toward an object located in the second area. Based on the actions of the identified person, the label set in the second area is changed to a different label from the label set based on the first machine learning model. Based on the video data including the aforementioned image data, tracking information obtained by tracking the same person is used to set reference lines indicating the person's movement path in the passage area of ​​the image data. Based on the skeletal information of each person captured in the aforementioned video data, the position of each person is determined. Using the position of each person, the movement trajectory of each person in the video data is identified. In the aforementioned image data, a plurality of clusters relating to the movement trajectory of each person are generated by clustering based on the distance between each reference line and the movement trajectory of each person. For each of the aforementioned clusters, a region of interest is extracted that includes clusters whose evaluation value, based on the angle between each movement trajectory belonging to the cluster and the baseline, is equal to or greater than a threshold. The labels set for each of the multiple areas configured based on the first machine learning model are changed based on the region of interest, including the cluster. A label changing program characterized by executing a process.

2. A computer, Acquire image data with multiple areas, By inputting the aforementioned image data into the first machine learning model, a label is set for each of the multiple areas. Identify the actions taken by a person located in the first of the aforementioned multiple areas toward an object located in the second area. Based on the actions of the identified person, the label set in the second area is changed to a different label from the label set based on the first machine learning model. From each image data within the video data, including the aforementioned image data, the position of each person appearing in the video data is identified. Based on the angle between the orientation of the person's face and the orientation of the person's body at each of the aforementioned person's positions, within the first area, a region of interest to be the subject of behavioral analysis of the person is identified. The labels of each of the multiple areas set based on the first machine learning model are changed based on the area of ​​interest. A label changing program characterized by executing a process.

3. The process for identifying the aforementioned action is: By inputting the acquired image data into a second machine learning model, skeletal information of the person located in the first area is generated. The process to be modified is: Based on the generated skeletal information, the actions taken by the person toward an object located in the second area are identified. The label changing program according to claim 1 or 2, characterized in that it performs a process of changing the label set in the second area using the identified action.

4. The process to be modified is: The aforementioned reference lines are set in the pathway regions identified by the first machine learning model, Multiple clusters relating to the image are generated by clustering based on the distance between each pixel belonging to the pathway region identified by the first machine learning model and each of the reference lines. Among the multiple clusters relating to the aforementioned image, identify the cluster of interest corresponding to the region of interest. The region of the aforementioned cluster of interest is modified to include the region of the corresponding region of interest. Based on the first machine learning model, the labels already set for the modified region of the cluster of interest are changed to labels corresponding to the region of interest. A label changing program according to claim 1, characterized by performing a process.

5. The process to be modified is: In the passage area identified by the first machine learning model, reference lines indicating the movement path of a person are set. Multiple clusters are generated by clustering based on the distance between each pixel belonging to the aforementioned passage region and each of the aforementioned reference lines. Among the multiple clusters mentioned above, identify the cluster of interest that corresponds to the area of ​​interest, The region of the aforementioned cluster of interest is modified to include the region of the corresponding region of interest. The label set for the modified region, which is set based on the first machine learning model, is changed to the label corresponding to the region of interest. The label changing program according to claim 2, characterized by performing a process.

6. Computers Acquire image data with multiple areas, By inputting the aforementioned image data into the first machine learning model, a label is set for each of the multiple areas. Identify the actions taken by a person located in the first of the aforementioned multiple areas toward an object located in the second area. Based on the actions of the identified person, the label set in the second area is changed to a different label from the label set based on the first machine learning model. Based on the video data including the aforementioned image data, tracking information obtained by tracking the same person is used to set reference lines indicating the person's movement path in the passage area of ​​the image data. Based on the skeletal information of each person captured in the aforementioned video data, the position of each person is determined. Using the position of each person, the movement trajectory of each person in the video data is identified. In the aforementioned image data, a plurality of clusters relating to the movement trajectory of each person are generated by clustering based on the distance between each reference line and the movement trajectory of each person. For each of the aforementioned clusters, a region of interest is extracted that includes clusters whose evaluation value, based on the angle between each movement trajectory belonging to the cluster and the baseline, is equal to or greater than a threshold. The labels set for each of the multiple areas configured based on the first machine learning model are changed based on the region of interest, including the cluster. A method for changing labels, characterized by executing a process.

7. Acquire image data with multiple areas, By inputting the aforementioned image data into the first machine learning model, a label is set for each of the multiple areas. Identify the actions taken by a person located in the first of the aforementioned multiple areas toward an object located in the second area. Based on the actions of the identified person, the label set in the second area is changed to a different label from the label set based on the first machine learning model. Based on the video data including the aforementioned image data, tracking information obtained by tracking the same person is used to set reference lines indicating the person's movement path in the passage area of ​​the image data. Based on the skeletal information of each person captured in the aforementioned video data, the position of each person is determined. Using the position of each person, the movement trajectory of each person in the video data is identified. In the aforementioned image data, a plurality of clusters relating to the movement trajectory of each person are generated by clustering based on the distance between each reference line and the movement trajectory of each person. For each of the aforementioned clusters, a region of interest is extracted that includes clusters whose evaluation value, based on the angle between each movement trajectory belonging to the cluster and the baseline, is equal to or greater than a threshold. The labels set for each of the multiple areas configured based on the first machine learning model are changed based on the region of interest, including the cluster. An information processing device characterized by having a control unit.

8. A computer, Acquire image data with multiple areas, By inputting the aforementioned image data into the first machine learning model, a label is set for each of the multiple areas. Identify the actions taken by a person located in the first of the aforementioned multiple areas toward an object located in the second area. Based on the actions of the identified person, the label set in the second area is changed to a different label from the label set based on the first machine learning model. From each image data within the video data, including the aforementioned image data, the position of each person appearing in the video data is identified. Based on the angle between the orientation of the person's face and the orientation of the person's body at each of the aforementioned person's positions, within the first area, a region of interest to be the subject of behavioral analysis of the person is identified. The labels of each of the multiple areas set based on the first machine learning model are changed based on the area of ​​interest. A method for changing labels, characterized by performing a process.

9. Obtain image data having multiple areas, By inputting the aforementioned image data into the first machine learning model, a label is set for each of the multiple areas. Identify the actions taken by a person located in the first of the aforementioned multiple areas toward an object located in the second area. Based on the actions of the identified person, the label set in the second area is changed to a different label from the label set based on the first machine learning model. From each image data within the video data, including the aforementioned image data, the position of each person appearing in the video data is identified. Based on the angle between the orientation of the person's face and the orientation of the person's body at each of the aforementioned person's positions, within the first area, a region of interest to be the subject of behavioral analysis of the person is identified. The labels of each of the multiple areas set based on the first machine learning model are changed based on the area of ​​interest. An information processing device characterized by having a control unit.