Information processing device and information processing method

The integration of video and 3D model data in the information processing system addresses visibility and accuracy issues in human traffic monitoring, offering detailed insights into customer behavior for marketing and operational improvements.

WO2025248742A1PCT designated stage Publication Date: 2025-12-04NTT DOCOMO INC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/019973
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-31
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing human traffic monitoring systems based solely on video footage from a single camera suffer from poor visibility and reduced accuracy due to limited angles of view and blind spots, making it impossible to effectively analyze customer behavior within a space.

Method used

An information processing system that combines video data from a camera with 3D model data from a 3D scanner to analyze human behavior by tracking the intersection of skeletal key points with specific surfaces in a virtual space, using a 3D model to enhance visibility and accuracy.

Benefits of technology

Provides highly convenient analysis results for human behavior, enabling insights into customer behavior and purchasing patterns, which can inform marketing strategies and improve store layouts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024019973_04122025_PF_FP_ABST
    Figure JP2024019973_04122025_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device according to one embodiment comprises: a first acquisition unit that acquires video data indicating a video obtained by imaging a certain real space; a second acquisition unit that acquires 3D model data of a virtual space indicating the real space; and an output unit that outputs behavioral data obtained by analyzing the trajectory of an intersection between a line obtained on the basis of a point included in 3D skeleton key points extracted from each of a plurality of persons appearing in the video and a specific surface of the virtual space, for at least some of the plurality of persons.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device and information processing method

[0001] The present invention relates to an information processing device and an information processing method.

[0002] Techniques for monitoring human traffic are known. For example, Patent Literature 1 discloses a system that uses a three-dimensional virtual space to efficiently monitor a target. In this system, a monitoring location is photographed with a camera. A change area is extracted from the photographed image. A three-dimensional virtual space corresponding to this space is set in advance, and the change area is converted into a plate-like object that simply represents a person and placed in this three-dimensional virtual space. A user monitors this three-dimensional virtual space.

[0003] Patent No. 5960472

[0004] The technology described in Patent Document 1 was not originally intended to analyze human behavior such as movement lines or staying positions.

[0005] In contrast, the present invention provides highly convenient analysis results based on improved techniques for analyzing human behavior.

[0006] One aspect of the present disclosure provides an information processing device having a first acquisition unit that acquires video data showing a video captured of a certain real space, a second acquisition unit that acquires 3D model data of a virtual space showing the real space, and an output unit that outputs behavioral data obtained by analyzing, for at least some of the multiple humans, the trajectory of the intersection between a line obtained based on points included in 3D skeletal key points extracted from each of the multiple humans shown in the video and a specific surface in the virtual space.

[0007] Another aspect of the present disclosure provides an information processing method including the steps of acquiring video data showing a video of a real space, acquiring 3D model data of a virtual space showing the real space, and outputting behavioral data obtained by analyzing, for at least some of the multiple humans appearing in the video, the trajectories of intersections between a line obtained based on points included in 3D skeletal key points extracted from each of the multiple humans appearing in the video and a specific surface in the virtual space.

[0008] According to the present invention, it is possible to present highly convenient analysis results based on improved techniques for analyzing human behavior.

[0009] 1 is a diagram illustrating an example of the system configuration of the information processing system 1. A diagram illustrating an example of the functional configuration of the information processing system 1. A diagram illustrating an example of the hardware configuration of an information processing device 10. A sequence chart illustrating an example of a method for acquiring 3D model data in the information processing system 1. A diagram illustrating an example of 3D model data of a virtual space V1. A diagram illustrating an example of a floor image of the virtual space V1. A diagram illustrating an example of a spatial database. A sequence chart illustrating an example of a method for aligning video data and 3D model data in the information processing system 1. A diagram illustrating an example of video data of a real space R1. A flowchart illustrating an example of a method for recording and outputting behavioral data in the information processing system 1. A diagram illustrating an example of a bone model. A diagram illustrating an example of behavioral data of an individual on floor surface S1. A diagram illustrating an example of behavioral data of multiple people on floor surface S1. A diagram illustrating an example of behavioral data of multiple people on surface S3. A diagram illustrating an example of behavioral data in a tabular format showing the aggregated results of people flow analysis.

[0010] 1. Configuration FIG. 1 is a diagram illustrating the system configuration of an information processing system 1. In this example, the information processing system 1 (or simply referred to as the system) is a system for analyzing behavior, such as the movement or presence of people in a target space (an example of a real space) based on a camera image of the target space. The target space and real space refer to any space that exists in the real world, such as various locations such as stores, facilities, roads, or outdoor spaces. In this example, the people flow analysis is performed by combining processing using a virtual space that recreates the real space on data. The virtual space refers to a three-dimensional (3D) virtual space constructed on a computer or a network.

[0011] However, when people flow analysis is performed solely based on video footage from a single camera installed in a target space, problems arise, such as poor visibility of the analysis results due to the camera's physically limited angle of view, and reduced accuracy of the people flow analysis due to potential blind spots in the angle of view. These problems cannot be avoided using only video footage of a store captured by a camera, making it impossible to perform sufficient analysis of customer behavior within the store, their interest in products, or their purchasing behavior. Therefore, this embodiment presents highly convenient analysis results based on improved technology for analyzing human behavior.

[0012] The information processing system 1 includes an information processing device 10, a camera 20, and a 3D scanner 30. The information processing device 10 is an information processing device or server device in the information processing system 1. In this example, the information processing device 10 acquires video data representing a video of a target space captured by the camera 20 and 3D model data of a virtual space representing the target space from the 3D scanner 30. Based on both the video data and the 3D model data, the information processing device 10 performs people flow analysis. Finally, the information processing device 10 outputs behavioral data analyzed for people captured in the video captured (or captured) by the camera 20. This person refers to a person to be analyzed for each target space. For example, if the target space is a store (such as a shop or a tenant), this would include customers visiting the store. Behavioral data refers to data analyzed and recorded from the perspectives of human behavior, particularly movement, stagnation, and movement. In addition to individual data, behavioral data includes data obtained by aggregating data from multiple people and processing (i.e., statistical processing) the data of multiple people according to specific criteria.

[0013] As an example, people flow analysis can be applied to marketing. Marketing applications include, for example, analyzing which areas of a store are most likely to attract customers, or how customers behave toward each product. Customer behavior includes information such as whether they make an immediate purchase decision or hesitate before purchasing. This information is useful for estimating a customer's latent purchasing psychology, which cannot be fully grasped by simply compiling data based on the customer's product purchase history. Such information is expected to contribute to the implementation of optimal store layouts, the improvement of service quality, and increased profits. System users (or simply referred to as "users") can use the results of this people flow analysis to formulate business strategies or operational policies.

[0014] The camera 20 is a photographing device in the information processing system 1 that photographs the target space. Examples of the camera 20 include a surveillance (security) camera, an analog camera, a network (IP) camera, a cloud camera, or an infrared camera. Typically, the camera 20 photographs the target space from a fixed position and a fixed angle of view. The camera 20 may be implemented with one or more units, a structure that allows the position to be moved or the angle of view to be adjusted, specialized photography of an arbitrary target, and / or an automatic detection (detection) function using AI or the like. The video captured by the camera 20 is transmitted to the information processing device 10 as video data as appropriate (periodically, periodically, or at any frequency). Here, "video" refers to video.

[0015] The 3D scanner 30 is a device in the information processing system 1 that generates 3D model data of a virtual space representing a real space. The 3D scanner 30 may include, for example, a smartphone equipped with LiDAR (Light Detection and Ranging), a handheld device, a laser device, a camera equipped with photogrammetry, or various sensors. The 3D scanner 30 scans the target space by a worker or a work robot. The scanning occurs during a preliminary stage (preparation stage) of people flow analysis by the system or when the state (layout) of the real space changes. Essentially, the 3D scanner 30 is not installed and functioning in the real space at all times, but is used by an operator as appropriate depending on the situation. The 3D model data generated by the 3D scanner 30 is transmitted to the information processing device 10.

[0016] 2 is a diagram illustrating an example of the functional configuration of the information processing system 1. In this embodiment, the information processing device 10 has functional blocks (components) including a first acquisition unit 11, a second acquisition unit 12, an output unit 13, a storage unit 18, and a control unit 19. In this example, the storage unit 18 stores various data and programs including a database, for example. In this example, the control unit 19 performs various controls.

[0017] The first acquisition unit 11 acquires video data showing a captured image of the target space. In this example, the first acquisition unit 11 acquires the video data in real time via a network, for example, from the camera 20. That is, the first acquisition unit 11 continuously acquires video data from the information processing device 10 and the camera 20.

[0018] The second acquisition unit 12 acquires 3D model data representing the target space. In this example, the first acquisition unit 11 acquires the 3D model data via a network, for example, from a 3D scanner 30. Note that a predetermined common file format is used for the 3D model data.

[0019] The output unit 13 outputs behavioral data obtained by analyzing the behavior of at least some of the multiple people shown in the video. Behavioral data obtained by analyzing human behavior refers to data obtained by aggregating behavior from at least one perspective or statistically processing behavior from at least one perspective. When multiple people are shown in the video, the behavioral data is data obtained by analyzing the behavior of at least some of the multiple people. "At least some of the people" includes cases where only one person is the target. When multiple people are shown in the video and the behavior of only some of the people is the subject of analysis, the person to be analyzed is selected by the system or the user. The behavioral data is obtained by analyzing, for example, aggregating or statistically processing, the trajectory of the intersection between a specific surface in the target section and the target person. The intersection is obtained based on a 3D model of the target space and a 3D model of the person. Specifically, the behavioral data includes information about the trajectory of the intersection between a line obtained based on points included in 3D skeletal keypoints extracted from the person shown in the video and a specific surface in the virtual space.

[0020] 3D skeletal keypoints are points that represent specific positions on the human body in 3D space. In this example, a human in a video is abstracted and then turned into a 3D model. An abstracted 3D model of a human is a 3D model that uses at least 3D skeletal keypoints extracted from the video and has been anonymized to prevent the identification of individuals from a privacy perspective. In one example, the human model is represented only by bones, which are the basic structure for moving the 3D model. The process of abstracting a human includes a bone conversion process that converts the human into a bone model, a process that converts specific body parts into arbitrary shapes, or a process that masks (hides) or reveals at least a portion of the human.

[0021] The intersection of a line obtained based on points included in the extracted 3D skeletal keypoints and a specific surface in virtual space is the intersection obtained when the human in the video is converted into the above-mentioned 3D model and this 3D model is placed in virtual space. The correspondence between positions in the video and positions in virtual space is set in advance. Note that, in this case, if the video shows only a portion of the human body rather than the entire body, the information processing device 10 estimates the 3D skeletal keypoints of the entire body from only the portion shown in the video. This is known as 3D model interpolation processing. For example, the information processing device 10 estimates the entire image of a human who is hidden behind an object from parts shown in the video (such as the position of the hands, the angle of the neck, or the direction of the gaze) and interpolates the 3D model.

[0022] The points included in the 3D skeletal keypoints refer to specific points (e.g., head and waist) selected from the 3D skeletal keypoints (e.g., head, shoulders, elbows, hands, waist, knees, and feet) extracted from each person appearing in the video. The line obtained based on the points included in the 3D skeletal keypoints is a line drawn based on the selected points. The selected points are specified by the user or automatically determined by the system. The information processing device 10 extracts 3D skeletal keypoints from the video and attaches labels indicating the body parts to the extracted points. The line drawn based on the selected points (hereinafter referred to as the "reference line") is, for example, a straight line passing through two of the 3D skeletal keypoints when these two points are selected. Specifically, the reference line is a straight line passing through the 3D skeletal keypoints of the person's head and waist, a straight line passing through the 3D skeletal keypoints of the elbows and hands, or a straight line passing through the 3D skeletal keypoints of the head and hands.

[0023] A specific plane in a virtual space, from a geometrical perspective, represents any two-dimensional plane in a three-dimensional space. More specifically, if the virtual space represents a store in real space, it refers to a plane including the store's floor, walls, and each surface of an object. In this application, the point where the reference line of the 3D model intersects with this plane is called an "intersection," and the result representing the movement of the intersection based on changes over time, etc., is called a "trajectory." For example, if the specific plane is the floor of the virtual space, the intersection corresponds to the position where the 3D model is standing on the floor.

[0024] The behavioral data may be data obtained by aggregating intersections that satisfy the condition of staying in a designated area on a specific plane of the virtual space for a threshold time or longer. The designated area refers to an area further inside any two-dimensional plane in the virtual space. This area can be designated in advance, for example, by an administrator using the system. In the above store example, this could be any area on the aisle side of the floor where product shelves are placed. In this case, it is possible to analyze how long a person stayed in front of a specific product shelf. The behavioral data includes the results of aggregating intersections that satisfy the condition by area.

[0025] The behavioral data may include image data of a map representing the results of statistical processing of intersection trajectories for multiple people, drawn on a specific surface in the virtual space. In one example, this map is a visualized graph, known as a "heat map," in which the number of intersections is represented by color or shade. The image data refers to image data captured from any direction by a virtual camera (modular) used in the virtual space, included in an application for processing 3D model data. The image is an orthoimage of the floor of the virtual space as seen from the ceiling of the virtual space. The virtual camera (or orthocamera) can capture images of the virtual space from specified directions, not just the ceiling, and is used to obtain image data including specific surfaces. This enables people flow analysis on various surfaces.

[0026] The targets of the virtual space include various objects in the target space. For example, consider a case where the virtual space has a 3D model corresponding to a non-architectural structure (e.g., so-called furniture such as a store shelf) installed in the target space, and a specific surface is a surface of the non-architectural structure exposed to an aisle. In one example, the intersection is a position on the surface of the 3D model of the non-architectural structure that corresponds to a position touched by a human hand in the image. In another example, the intersection corresponds to a position on this surface where a human's line of sight intersects in the image.

[0027] FIG. 3 is a diagram illustrating an example of the hardware configuration of the information processing device 10. The information processing device 10 is physically configured as a computer including a processor 101, a memory 102, a storage 103, a communication device 104, an input device (optional), a display device (optional), and a bus connecting these. Each of these devices operates using power supplied from a battery (not shown). In the following description, the term "device" can be interpreted as a circuit, device, unit, etc. The hardware configuration of the information processing device 10 may be configured to include one or more of the devices shown in FIG. 3, or may be configured without including some of the devices. Furthermore, the information processing device 10 may be configured by communicating and connecting multiple devices each having a different housing.

[0028] Each function of the information processing device 10 is realized by loading specified software (programs) onto hardware such as the processor 101, memory 102, etc., so that the processor 101 performs calculations, controls communication via the communication device 104, and controls at least one of reading and writing data in the memory 102 and storage 103.

[0029] The processor 101 controls the entire computer by running, for example, an operating system. The processor 101 may be configured as a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, a register, etc. Furthermore, for example, a baseband signal processing unit, a call processing unit, etc. may be realized by the processor 101.

[0030] The processor 101 reads programs (program codes), software modules, data, etc. from at least one of the storage 103 and the communication device 104 into the memory 102 and executes various processes in accordance with these. The programs used are those that cause a computer to execute at least some of the operations described below. The functional blocks of the information processing device 10 may be implemented by a control program stored in the memory 102 and running on the processor 101. Various processes may be executed by one processor 101, or may be executed simultaneously or sequentially by two or more processors 101. The processor 101 may be implemented by one or more chips. The programs may be transmitted to the information processing device 10 via a telecommunications line.

[0031] The memory 102 is a computer-readable recording medium and may be configured by, for example, at least one of a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a random access memory (RAM), etc. The memory 102 may also be called a register, a cache, a main memory (primary storage device), etc. The memory 102 can store executable programs (program codes), software modules, etc. for implementing the method according to this embodiment.

[0032] Storage 103 is a computer-readable recording medium, and may be composed of at least one of, for example, an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disk), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, etc. Storage 103 may also be called an auxiliary storage device.

[0033] The communication device 104 is hardware (transmission / reception device) for communicating between computers via at least one of a wired network and a wireless network, and is also called, for example, a network device, a network controller, a network card, or a communication module.

[0034] Each device, such as the processor 101 and the memory 102, is connected by a bus for communicating information. The bus may be configured using a single bus, or different buses may be used between each device.

[0035] The information processing device 10 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 101 may be implemented using at least one of these pieces of hardware.

[0036] In this example, the programs stored in the storage 103 include a program for causing a computer to function as a server of the information processing system 1 (hereinafter referred to as a “server program”).

[0037] When the processor 101 is executing a server program, the processor 101, memory 102, storage 103, and communication device 104 are examples of functional blocks for operating the information processing device 10. The processor 101 is an example of a control unit 19. At least one of the memory 102 and the storage 103 is an example of a storage unit 18. The communication device 104 is an example of a first acquisition unit 11, a second acquisition unit 12, and an output unit 13. The configuration of the information processing system 1 has been described above. Next, the operation of the information processing system 1 will be described.

[0038] 4 is a sequence chart illustrating a method for acquiring 3D model data in the information processing system 1. Here, a process will be described in which the information processing device 10 acquires 3D model data of a target space as a preparatory step before performing processing to analyze a person's movement line or presence in the system.

[0039] In step S101, the 3D scanner 30 scans the target space for people flow analysis using various sensors. The scope of the real space to be scanned can be freely determined based on the results of actual scans performed by an operator or other personnel. The target space can represent not only architectural structures such as walls, floors, and ceilings, but also non-architectural structures such as furniture and electrical appliances present within the target space. Therefore, for example, if the location to be scanned is a store (both inside and outside), the 3D model of the target space will include not only the architectural structural objects that form the store space, but also non-architectural structural objects such as product shelves installed within the store. The level of detail with which these objects are represented can be set according to the specifications required for the system. For example, a product shelf object may be a simple rectangular object without reproducing the unevenness of each shelf.

[0040] In step S102, the information processing device 10 acquires 3D model data representing the target space. The 3D scanner 30 scans the target space and simultaneously generates 3D model data of the target space and records the data in a database or the like. Once generation of the 3D model data is complete, the 3D scanner 30 transmits the data to the information processing device 10 via a network. Here, the 3D model data (i.e., virtual space) will be described.

[0041] FIG. 5 is a diagram illustrating 3D model data of a virtual space V1. In this example, the 3D model data of the virtual space V1 is obtained by scanning the real space R1 with a 3D scanner 30. The virtual space V1, for example, 3D-reproduces objects such as the floor, walls, or shelves of the real space R1. In FIG. 5, individual products (items) on the shelves are not represented as objects, but photos of the products displayed on the shelves are attached to the shelf objects as textures. Thus, the virtual space V1 includes a shelf object with specific products displayed on the shelves. Depending on the resolution of the images attached as textures, it is possible to reproduce even details such as the shape of each product or illustrations or text printed on the packaging. Furthermore, the virtual space V1 can be viewed via a display by a user operating the system using dedicated software. While FIG. 5 merely shows a perspective view of the 3D model data from a certain viewpoint, for example, a user can observe the structure of the virtual space V1 from various viewpoints using the function of a virtual camera VC1 incorporated into the software. The virtual camera VC1 will be described later.

[0042] Returning to FIG. 4 , in step S103, the information processing device 10 determines a target surface (hereinafter referred to as a "target surface"; an example of a specific surface) among the surfaces included in the virtual space in the acquired 3D model data. A user (e.g., a store manager) specifies which surface among the surfaces included in the virtual space is to be the target surface. Here, an example will be described in which the floor surface is specified as the target surface. For example, the information processing device 10 calculates a floor equation for the 3D model data and acquires the coordinates of the center of gravity of the 3D model (virtual space) based on the calculated floor equation. The information processing device 10 then adjusts the angle (angle of view) of the virtual camera so that it is orthogonal to the floor equation (perpendicular to the floor), and installs the virtual camera at a position that overlaps with the acquired coordinates of the center of gravity. The information processing device 10 sets the virtual camera to ortho mode, captures an image, and acquires a floor image.

[0043] 5, the information processing device 10 takes images using a virtual camera VC1 installed on the ceiling (the positional relationship is approximate). When a system administrator or the like intervenes in the system, the virtual camera VC1 can be operated to take images of the surface of the virtual space V from various directions. Here, a simplified diagram of a floor image will be used to explain the image as an example.

[0044] 6 is a diagram illustrating an example of a floor image of a virtual space V1. In this example, floor image I1 represents an image obtained when the virtual space V1 is photographed in ortho mode from a viewpoint from the ceiling of a virtual camera VC1 (arrow VA1 in FIG. 5). Floor image I1 is an image of a surface S1, which is the floor, viewed from a viewpoint in the vertical direction. Surface S1 is an example of a target surface.

[0045] Returning to Fig. 4, in step S104, the information processing apparatus 10 sets a target area (an example of a region) for the acquired plane image. The target area can be designated by, for example, the user.

[0046] Referring again to FIG. 6 , in the example of FIG. 6 , area A11, which is defined from the left on floor surface S1 on the aisle side in front of refrigerator Rf1, and area A12, which is defined on floor surface S1 on the aisle side in front of center shelf Sh1, are areas designated by the user as target areas. Conditions related to data aggregation are set in advance for these areas. These conditions are set, for example, by the user. These conditions include conditions related to stay time, such as a threshold value. Details of the areas will be described later.

[0047] As described above, the information processing device 10 can set specific surfaces and areas in advance for the 3D model data acquired from the 3D scanner 30. These settings are recorded in a database or the like. Here, the database for managing 3D model data will be described.

[0048] FIG. 7 is a diagram illustrating a spatial database. In this example, the spatial database 1001 includes multiple records. Each record corresponds to a record for each real space and each virtual space. Each record includes a real space ID, a virtual space ID, a surface ID, surface information, an area ID, and area information. The real space ID and virtual space ID are identifiable information uniquely assigned to the real space from which the scan was performed and the virtual space representing that real space in 3D model data. In principle, one real space (e.g., real space R1) corresponds to one virtual space (e.g., virtual space V1) generated by scanning. The surface ID and surface information are unique identification information and surface information representing a specific surface in the virtual space. For example, if the surface information represents a floor surface, it represents surface S1 as viewed from the ceiling of virtual space V1 in FIG. 6 by virtual camera VC1 (ortho camera). The area ID and area information are unique identification information and area information representing a specific area divided (specified) as an area for each surface. For example, when the area information indicates "in front of refrigerator Rf1," this refers to area A11 on the floor surface in FIG. 6. In this example, the surface information and area information may include, for example, coordinate information (position information relative to the virtual space). In addition, specifications, thresholds, conditions, and the like determined by an administrator or the like may be recorded in advance in spatial database 1001 together with various data.

[0049] Next, a method for acquiring video data by the camera 20 and a method for aligning the video data with the 3D model data by the information processing device 10 will be described.

[0050] 2-2. Alignment Method Between Video Data and 3D Model Data FIG. 8 is a sequence chart illustrating an example of an alignment method between video data and 3D model data in the information processing system 1. Here, we will explain the process in which the information processing device 10 acquires video data of the target space as a preparatory step before analysis, and aligns the video data with the 3D model data acquired by the information processing device 10 in Section 2-1. In step S201, the camera 20 begins capturing images of the target space. In this case, the camera 20 installed on-site (or newly prepared) is usually installed in a fixed position, and the angle of view, i.e., the direction of the camera lens, is fixed. The camera 20 installed in this state can capture images of the target space and record the images.

[0051] In step S202, the information processing device 10 acquires video data showing an image of the target space captured by the camera 20. The information processing device 10 and the camera 20 are connected via the Internet or the like. Therefore, the camera 20 can continuously transmit the video data to the information processing device 10 via the network. Here, the video data will be described.

[0052] FIG. 9 is a diagram illustrating video data of real space R1. In this example, the video data is obtained by camera 20 capturing an image of real space R1. The video data is composed of, for example, multiple frames (images captured per unit time). In FIG. 9, video screen P1 represents one frame of the video data. Note that customer C1 represents a person appearing in the image. When the information processing device 10 acquires the video data from camera 20, it records it in a database. This may be the same database as spatial database 1001 or a database in which a correspondence relationship with the target space is recorded.

[0053] Returning to Fig. 8 , in step S203, the information processing device 10 reads 3D model data of the virtual space from the space database 1001. For example, if the real space R1 is shown in the video data or if the identification information of the video data acquired from the camera 20 indicates the real space R1, the information processing device 10 acquires data of the virtual space V1 corresponding to the real space R1 from the space database 1001.

[0054] In step S204, the information processing device 10 aligns the acquired video data with the read 3D model data. Specifically, the information processing device 10 pre-records in a database or the like the correspondence between feature points in the real space R1 shown in the video and their positions in the virtual space V1. This involves, for example, determining parameters that determine the correspondence between feature points in the real space R1 and feature points in the virtual space V1 and recording these parameters. Feature points in the virtual space include, for example, intersections between walls, intersections between walls and floors, intersections between walls and ceilings, and intersections between shelves and floors. In FIG. 9 , the information processing device 10 may pre-record the positional relationship of the floor or shelf shown on the video screen P1, as well as the corresponding position in the virtual space V1 of the location where customer C1 is shown, and may acquire an algorithm capable of detecting a human position using machine learning or the like.

[0055] In addition, when multiple cameras 20 exist in one target space (the system has multiple cameras 20), the information processing device 10 acquires video data from each of the multiple cameras 20, integrates the video data, and records it as a parallel combined data set. This enables the system to perform people flow analysis from multiple angles. Furthermore, the processing of steps S203 to S204 may be performed as appropriate in response to updates to the 3D model data. As described above, the information processing system 1 completes advance preparations through the processing described in sections 2-1 and 2-2. Next, a specific processing flow for people flow analysis will be described.

[0056] 2-3. Behavioral Data Recording Method The following describes operations related to recording and outputting behavioral data in the information processing system 1. Here, we will explain an example in which, after the advance preparations in sections 2-1 and 2-2, the behavior of actual people in the target space is recorded and analyzed.

[0057] FIG. 10 is a flowchart illustrating a method for recording and outputting behavioral data in the information processing system 1. The following operations particularly represent processing executed primarily by the information processing device 10. In step S1, the information processing device 10 acquires video data from the camera 20. Specifically, the camera 20 first captures an image of the target space. The camera 20 periodically transmits the video data to the information processing device 10 via various networks. Below, we will explain the processing of the information processing device 10 on, for example, one frame of video (one image) from the acquired video data.

[0058] In step S2, the information processing device 10 detects people appearing in the video. If there are multiple people, the information processing device 10 uniquely identifies and detects the multiple people. The information processing device 10 assigns an identifier to the extracted people. For people who have continued to exist since the previous frame, the information processing device 10 continues to assign the same identifier. Whether the extracted person is the same person in the previous frame and the current frame is determined based on their position and appearance characteristics.

[0059] In step S3, the information processing device 10 extracts 3D skeletal keypoints from the human detected in the video. Here, the extracted 3D skeletal keypoints will be explained using a bone model. A bone model is a model that represents a human using only bones, which are the basic structure for moving a 3D model. The information processing device 10 does not need to generate a 3D model of the human in the following process, but only needs to calculate the required coordinates using the extracted skeletal keypoints. However, for ease of understanding, the explanation will be given here of creating a 3D model of a human.

[0060] FIG. 11 is a diagram illustrating a bone model. In this example, two examples of the process for generating a bone model are introduced. The first example is an example in which a person appearing in a video is detected and a bone generation process (an example of abstraction processing) is performed. Generally, skin mesh models are used in the field of 3D models, which are models in which meshes are added to pre-set bones. Here, skin and mesh are not used, and the information processing device 10 replaces the person in the video with bones. In this example, customer C11 represents a person appearing in the video. Although shown as a silhouette for simplicity, customer C11 maintains a posture facing forward relative to the video screen P11. Here, the information processing device 10 extracts 3D skeletal key points from customer C11 appearing in the video. The information processing device 10 generates bones from the extracted 3D skeletal key points (this process is referred to as "bone generation processing") to obtain a bone model B11. The bone model B11 is formed from multiple bones. Although details are omitted, this three-dimensional structure plays the role of a joint in the human body. The shape of the bone may be any shape, such as a line, a plane, or a solid. The information processing device 10 records information that identifies the acquired bone model B11. The information that identifies the bone model includes, for example, information indicating the coordinates of each bone.

[0061] The second example is a case in which bone generation processing is performed when at least a portion of a person in the video is obscured by an obstruction or the like and is not visible to the camera 20 (in the case of Figure 11, the person is not displayed on the video screen P12). In this example, customer C12 represents a portion of the person in the video (part of the face, part of the hand, part of the arm). Here, the information processing device 10 first detects customer C12 as a person. Next, based on the portion of the body part displayed on the video screen P12, the information processing device 10 estimates the image of the person, that is, how the person would appear in the video if they were actually present, and performs bone generation processing. Bone model B12 represents the interpolated model obtained by the information processing device 10 through the above processing. For example, a bone model B12 with one hand raised (and facing forward) is generated based on the position or angle of the person's hand H12 displayed on the video screen P12. In this way, by processing the people in the video data, the system can protect the privacy of each individual person without identifying them. Therefore, at the same time as generating the bone model, the system may strive to protect privacy by deleting, masking, or encrypting images of humans in the video data.

[0062] Returning to Fig. 10, in step S4, the information processing device 10 estimates the position in the virtual space of the detected human (or bone model). Specifically, the information processing device 10 estimates the position in the virtual space based on the video data from the positional relationships of objects, floors, shelves, products, other people, etc., relative to the human in the video. Note that the position estimation is performed based on alignment between the video data and 3D model data.

[0063] Here, the processing of steps S3 and S4 may be optimized as appropriate, for example, the order of the steps may be reversed, or the steps themselves may be combined. In other words, in one example, the information processing device 10 may correct or abstract the detected human when placing the human at a corresponding position in the virtual space.

[0064] In step S5, the information processing device 10 calculates the position of the intersection between the reference line obtained from the bone model and a predetermined specific surface in the virtual space. If there are multiple intersections or if the intersections are successively formed into a straight line or a curve, information corresponding to the calculation result according to the shape of the intersection (or the intersection line) with the surface may be acquired.

[0065] In step S6, the information processing device 10 records the trajectory of the intersections. For example, the information processing device 10 can repeatedly perform the processes of steps S2 to S5 described above and record the change in the position of the intersections in multiple temporally consecutive frames as a trajectory. The information processing device 10 associates each intersection with an identifier of the 3D model and records it in the database. If there are multiple people in the frame, the information processing device 10 records the trajectory of the intersections for each person. In this way, it is possible to record the behavior data of people included in multiple frames that make up the video data.

[0066] In step S7, the information processing device 10 determines whether an output condition for the behavioral data is satisfied. The output condition is a preset condition, such as receipt of an explicit output instruction from a system user, arrival of a time specified by the system (e.g., midnight every day), or when the amount of data exceeds a threshold value.

[0067] If the output condition is not satisfied (step S7: NO), the information processing device 10 returns the process to step S1 (loops the process). In this example, the processes of steps S1 to S6 are repeatedly executed until the output condition is satisfied. Note that the information processing device 10 and the camera 20 can share video data over a continuous period of time via a network. Therefore, the information processing device 10 repeats the above process each time it acquires video data transmitted from the camera 20 until the output condition is satisfied. Note that the information processing device 10 does not need to perform the processes of steps S1 to S6 for all frames of the video, and may perform these processes every predetermined number of frames (for example, once every 30 frames).

[0068] On the other hand, if the output condition is satisfied (step S7: YES), the information processing device 10 proceeds to step S8.

[0069] In step S8, the information processing device 10 outputs behavioral data. The behavioral data includes, for example, data on the trajectories of intersections, which are aggregated (e.g., statistically processed) by any combination of items, such as for each individual, for each group of people, for each surface, for each area, or for each time period. The data types (or file formats) of the output behavioral data include, for example, image data that visualizes the trajectories of intersections on a surface, image data of a heat map that represents the density of intersections on a surface using shades of gray, or comma-separated variables (CSV) data that converts information about the trajectories of intersections into text and aggregates it. Here, several specific examples of behavioral data will be described.

[0070] FIG. 12 illustrates an example of individual behavior data on floor surface S1. In this example, image I11 is an image of floor surface S1 captured by capturing virtual space V1 in ortho mode. Intersection X1 is one of the intersections included in the behavior data of a specific individual. Intersections are represented by dots in the figure. The intersections trace a trajectory as shown in the figure as they change over time. Because intersections are recorded per unit time, the density of intersections increases in areas where customers spend a long time. If intersections are densely concentrated in a specific area (e.g., area A12 in front of center shelf Sh1 or area A13 in front of refrigerator Rf2) (i.e., the customer spent a long time in a specific area or went back and forth), it is possible to estimate, for example, which shelf and which product the customer is interested in. In this example, image I12 is an image in which the number (density) of intersections is depicted as a heat map based on the individual's trajectory of intersections under the same conditions as image I11. In the heat map, darker areas indicate areas where customers spend a long time (where intersections are densely packed). Conversely, lighter areas indicate areas where customers have not set foot or where they have spent a short time. The heat map drawn in image I12 can provide more convenient analysis results. Note that areas may or may not be displayed in the image. Furthermore, in addition to the image or heat map showing the trajectories of these intersections, the information processing device 10 may simultaneously output and record CSV data that lists the coordinates, time, surface, or area of ​​the intersections and is converted into text.

[0071] As a further variation of behavioral data, we present below images and heat maps of statistical processing performed on multiple people on the trajectories of intersections with specific surfaces.

[0072] FIG. 13 illustrates behavioral data of multiple people on floor surface S1. In this example, image I21 represents an image of floor surface S1 in virtual space V1. Image I21 includes trajectories of intersections for multiple people. Image I22 is an image depicting the trajectories of intersections as a heat map. Based on these images, it is possible to analyze the behavior of multiple people. For example, system users (store managers) can obtain insights that can be used for marketing, such as: people's attention is drawn to a particular product on a particular shelf; this aisle is a commonly used route, so placing a new product there will attract attention; or customers spend too much time waiting in front of the register, so automated registers should be introduced. Note that areas may or may not be displayed in the image, and intersections within and outside the areas may be arbitrarily processed and depicted on the image.

[0073] FIG. 14 is a diagram illustrating behavioral data of multiple people on surface S3. In this example, surface S3 corresponds to the surface of refrigerators Rf1-Rf3 exposed to the aisle (i.e., the view of the products in the refrigerators from the front, as seen from the customer's perspective). Surfaces include the surfaces of spaces and objects such as shelves, refrigerators, or walls. For example, the image depicts a locus of intersections where parts of a 3D human model corresponding to the hands or eyes (gaze) intersect with the surface. Image I31 is an image showing loci of intersections for multiple people on surface S3. Image I32 is an image depicting the loci of intersections as a heat map. For example, various regions are set on surface S3 within each image. These regions are set, for example, for each refrigerator, shelf, or product. In particular, areas A31 and A32 are set for each type of product, and the information processing device 10 can output statistically processed behavioral data for each region in various formats. Based on these images, it becomes possible to analyze the behavior of multiple people. Behavioral data such as product placement that makes it easy (or difficult) to take out of the refrigerator, popular (or unpopular) products, or sales of new products are very useful for analyzing customer purchasing behavior. Note that in these images, areas may or may not be displayed.

[0074] The information processing device 10 may output multiple types of behavioral data as described above, or may output only the type of behavioral data specified by the user. As a result, the information processing device 10 can present highly convenient analysis results based on improved techniques for analyzing human behavior.

[0075] In steps S1 to S3, when the information processing device 10 acquires video data from the camera 20 in real time, the information processing device 10 may output the video to a display device (such as an external monitor for monitoring) connected to the device. Alternatively, the information processing device 10 may record the video as video data. At this time, the information processing device 10 detects people appearing in the video and performs abstraction processing. The people appearing in the video are replaced with abstracted models. When the video is output to a display device, system users can observe customer behavior in real time while ensuring privacy protection. Furthermore, when the video data is recorded, the information processing device 10 can collect or store analyzable data without violating privacy protection standards.

[0076] 3. Modifications The present invention is not limited to the above-described embodiment, and various modifications are possible. Some modifications will be described below. Two or more of the following features may be combined and applied.

[0077] (1) Information Processing System 1 The hardware configuration and network configuration of the information processing system 1 are not limited to those exemplified in the embodiment. The information processing system 1 may have any hardware configuration and network configuration as long as the required functions can be realized. For example, multiple physical devices may cooperate to function as the information processing system 1. For example, the information processing system 1 may include a terminal (hereinafter referred to as an "administrator terminal or user terminal") held by an administrator who manages the target space. The administrator terminal may acquire behavioral data from the information processing device 10. Conversely, the information processing device 10 may acquire video data or 3D model data via the administrator terminal.

[0078] (2) Information Processing Device 10 The correspondence between the functional elements of the information processing device 10 and the hardware is not limited to that exemplified in the embodiment. For example, in the embodiment, at least some of the functions described as being implemented in the information processing device 10 may be implemented in another device or system, and conversely, at least some of the functions described as being implemented in another device or system may be implemented in the information processing device 10.

[0079] (3) Camera 20 The camera 20 is not limited to the one exemplified in the embodiment. The camera 20 may have any hardware configuration as long as it can realize the required functions and operations. The camera 20 may be equipped with a thermography camera. The camera 20 may also be equipped with a motion camera (sensor). This makes it possible to detect the state of the customer's movements. The camera 20 may also be equipped with a sound collection device such as a microphone.

[0080] Furthermore, in the embodiment, an example has been described in which the information processing device 10 continuously acquires and processes video from the camera 20 in real time. However, the method by which the information processing device 10 acquires video data representing video captured of the target space (i.e., the method by which the first acquisition unit 11 acquires video data) is not limited to the example described in the embodiment. The information processing device 10 may acquire video data that has been captured and accumulated in advance, rather than in real time. Video data captured in advance by the camera 20 may be accumulated in another device or the camera 20 itself, and uploaded to the information processing device 10 at a later date. Alternatively, video data captured in advance by the camera 20 may be sent to the information processing device 10 in real time, but the information processing device 10 may accumulate the video data without processing it in real time, and then perform the processes of steps S2 to S8 on the accumulated video data when a triggering event occurs.

[0081] (4) 3D Scanner 30 The 3D scanner 30 is not limited to the one exemplified in the embodiment. The 3D scanner 30 may have any hardware configuration as long as it can realize the required functions and operations. The 3D scanner 30 may be implemented in a robot or the like. When acquiring 3D model data of a real space, a robot equipped with the 3D scanner 30 may move around the target space and perform scanning on behalf of an administrator or worker.

[0082] (5) Method for Acquiring 3D Model Data The sequence chart shown in FIG. 4 merely illustrates one example of operations, and the method for acquiring 3D model data in the information processing system 1 is not limited to this. Some of the illustrated operations may be changed or omitted, the order may be changed, or new operations may be added. In step S103, when the information processing device 10 receives a designation of a specific surface in the virtual space from the user, the information processing device 10 may predefine a specific part of the 3D model that is determined to be an intersection with the surface. The information processing device 10 records the correspondence between the specific surface (e.g., the floor) and a specific part of the 3D model (e.g., a foot) in a database. Note that the processing in steps S103 and S104 may be performed after aligning the video data as described in Section 2-2.

[0083] (6) Alignment Method Between Video Data and 3D Model Data The sequence chart shown in Figure 8 merely shows one example of the operations, and the method for aligning video data and 3D model data in the information processing system 1 is not limited to this. Some of the illustrated operations may be changed or omitted, the order may be changed, or new operations may be added. In step S204, the information processing device 10 aligns the acquired video data with the read-out 3D model data, and then may adjust, correct, or change the position in the virtual space corresponding to the position where the person is shown, based on changes over time in the video data actually acquired from the camera 20.

[0084] (7) Behavioral Data Recording and Output Method The flowchart shown in FIG. 10 merely illustrates an example of the operation, and the method of recording and outputting behavioral data in the information processing system 1 is not limited thereto. Some of the illustrated operations may be changed or omitted, the order may be changed, or new operations may be added. The processing in steps S2 to S5 may be performed on multiple frames of video data rather than on a single frame of image. In step S8, the information processing device 10 may perform various analyses based on the trajectory of the intersections and output behavioral data. For example, if the surface represents a floor surface S1, a predetermined area may be defined as a "stopping area," and the number of people who have stayed within the intersection area for a threshold time or more (e.g., 5 seconds or more) may be tallied for each area. Alternatively, if the surface represents a surface S3 exposed to the aisle, such as a shelf (a surface located in a direction where products are visible to customers), an area defined for each product and where a specific part of the intersection is conditioned as "hand" may be defined as a "hand touch area," and the number of people who have stayed within the area for a threshold time or more (e.g., 2 seconds or more) may be tallied for each product. Alternatively, the locations where touches were made (such as the bottle body, label, or edge of the bottle) may be counted individually, or the number of times touches were made may be counted. These counting results may be recorded and output as tabular data (e.g., CSV).

[0085] FIG. 15 is a diagram illustrating an example of behavioral data in a tabular format showing the aggregated results of people flow analysis. The behavioral data 2001 is generated by aggregating the system's analysis results in the real space R1, which is the target space, and recording and outputting them as tabular text data (CSV). In this example, the behavioral data 2001 is composed of multiple records. Each record corresponds to a record for each real space (or each virtual space). Each record includes a real space ID, a target period, a surface ID or surface information, aggregated information, an area ID or area information, and a number of people. The real space ID, surface ID (surface information), and area ID (area information) are ID information or identification information shared with the spatial database 1001 described above. The target period is the period (time period) for which people flow analysis is performed. The target period can be freely set by system users or others when outputting data. The aggregated information is an aggregate item for aggregating people who meet conditions defined for each area, such as "stopping" or "hand touching." The aggregated information is predetermined by the system. Alternatively, the system user or others specify the conditions in advance. The number of people is the total number of people corresponding to the aggregated information during the target period. For example, the behavioral data 2001 may output a result indicating that, among customers who visited the store during a specific target period, 30 customers "stopped" (i.e., stayed for more than five seconds) in a certain area A11 (in front of refrigerator Rf1). Alternatively, the output may indicate that, during the same period, 20 customers "touched" (i.e., touched for more than two seconds) in area A31 (XX Brown), among the areas set for each product in the refrigerator. These aggregated results are expected to be applicable to marketing, and may be effective in psychological analysis of customer behavior, improvements to store layout, or improvements to the quality of product (service) provision.

[0086] Furthermore, the human abstraction process used to calculate behavioral data is not limited to bone modeling. Any process using 3D skeletal keypoints extracted from a human captured in a video may be used. For example, the information processing device 10 may simply extract 3D skeletal keypoints without converting the human captured in a video into a 3D model. In this case, the information processing device 10 may extract only a selected portion of the 3D skeletal keypoints rather than extracting all 3D skeletal keypoints that can be extracted from the human captured in the video. In this case, the information processing device 10 may extract only selected 3D skeletal keypoints from the video, determine the equation of the reference line without converting these into a 3D model, and calculate the intersection point with a specific surface of the 3D model in the target space. Alternatively, a skin mesh model that does not identify individuals and has the same appearance for all humans, for example, may be used instead of a bone model. Alternatively, a model in a fixed pose with immobile arms or legs may be used.

[0087] (8) Database The database (or the data itself) of the information processing system 1 shown in FIGS. 7 and 15 is not limited to the example shown in the embodiment. In this example, any type of data may be registered in the database. For example, the data recorded in the spatial database 1001 may include coordinate information of a virtual space, a surface, or an area. In this case, for example, the virtual space V1 is defined by three-dimensional coordinates, and the surface and the area are defined by a surface equation representing a two-dimensional plane. The range of the area may also be defined by coordinates or the like. In addition, a specific portion of the 3D model for determining an intersection for each surface (each area) may be recorded in advance in the database.

[0088] (9) Behavioral Data The behavioral data is not limited to those exemplified in the embodiments. The behavioral data may be data indicating the behavior of a single individual, or may be data obtained by statistically processing the behavior of multiple people. The behavioral data may be data processed with or without identifying a person. The behavioral data may be data targeting objects or events other than people. The behavioral data may be in any file format. The information processing device 10 may output the behavioral data to various devices depending on the purpose.

[0089] (10) Others The various programs executed by the processor 101 may be provided by downloading via a network such as the Internet, or may be provided in a state recorded on a computer-readable non-transitory recording medium such as a DVD-ROM. Each processor may be, for example, a CPU, an MPU (Micro Processing Unit), or a GPU (Graphics Processing Unit).

[0090] The block diagrams used to explain the above embodiments show functional blocks. These functional blocks (components) are realized by any combination of hardware and / or software. Furthermore, the method for realizing each functional block is not particularly limited. That is, each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are directly or indirectly connected (e.g., wired, wireless, etc.) and these multiple devices. The functional block may also be realized by combining software with the single device or multiple devices.

[0091] Functions include, but are not limited to, judgment, determination, judgment, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, consideration, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assignment. For example, a functional block (component) that performs transmission is called a transmitting unit or transmitter. As mentioned above, there are no particular limitations on how these functions are implemented.

[0092] For example, the information processing device 10 according to an embodiment of the present disclosure may function as a computer that performs the processing of the present disclosure.

[0093] Each aspect / embodiment described in the present disclosure may be applied to at least one of systems using LTE (Long Term Evolution), LTE-Advanced (LTE-A), SUPER 3G, IMT-Advanced, 4G (4th generation mobile communication system), 5G (5th generation mobile communication system), FRA (Future Radio Access), NR (New Radio), W-CDMA (registered trademark), GSM (registered trademark), CDMA2000, UMB (Ultra Mobile Broadband), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, UWB (Ultra-Wideband), Bluetooth (registered trademark), or other suitable systems, and next-generation systems enhanced based on these. Furthermore, a combination of multiple systems (e.g., a combination of at least one of LTE and LTE-A with 5G) may also be applied.

[0094] The order of the procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.

[0095] Input and output information may be stored in a specific location (for example, memory) or may be managed using a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be sent to another device.

[0096] The determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).

[0097] Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure.

[0098] Software, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise, should be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc. Additionally, software, instructions, information, etc. may be transmitted or received over a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave), then such wired and / or wireless technologies are included within the definition of a transmission medium.

[0099] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof. Note that terms described in this disclosure and terms necessary for understanding this disclosure may be replaced with terms having the same or similar meanings.

[0100] Furthermore, the information, parameters, etc. described in this disclosure may be expressed using absolute values, may be expressed using relative values ​​from a predetermined value, or may be expressed using other corresponding information.

[0101] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly stated otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."

[0102] As used in this disclosure, any reference to an element using a designation such as "first," "second," etc. does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must in some way precede the second element.

[0103] The "unit" in the configuration of each of the above devices may be replaced with "means," "circuit," "device," or the like.

[0104] When the terms "include," "including," and variations thereof are used in this disclosure, these terms are intended to be inclusive, similar to the term "comprising." Furthermore, when the term "or" is used in this disclosure, it is not intended to be an exclusive or.

[0105] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.

[0106] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different."

[0107] 1...information processing system, 10...information processing device, 20...camera, 30...3D scanner, 11...first acquisition unit, 12...second acquisition unit, 13...output unit, 18...storage unit, 19...control unit, 101...processor, 102...memory, 103...storage, 104...communication device, 1001...spatial database, 2001...behavioral data

Claims

1. An information processing device having: a first acquisition unit that acquires video data showing a video of a real space; a second acquisition unit that acquires 3D model data of a virtual space showing said real space; and an output unit that outputs behavioral data obtained by analyzing, for at least some of the plurality of people, the trajectory of the intersection between a reference line obtained based on points included in 3D skeletal key points extracted from each of said plurality of people shown in said video and a specific surface of said virtual space.

2. The information processing device according to claim 1, wherein the behavioral data is data obtained by aggregating intersections that satisfy the condition that the user has stayed in a specified area on the specific surface of the virtual space for a threshold time or longer.

3. The information processing device according to claim 2, wherein the behavioral data includes a result of tallying intersections that satisfy the condition for each of the areas.

4. The information processing device according to claim 1, wherein the output unit outputs data obtained by statistically processing the trajectories for at least some of the people as the behavioral data.

5. The information processing device according to claim 4, wherein the behavioral data includes image data of an image in which a map obtained by statistically processing the trajectories of the intersections for at least some of the people is drawn on the specific surface of the virtual space.

6. The information processing device according to claim 5, wherein the image is an orthoimage of the floor of the virtual space as viewed from the ceiling of the virtual space.

7. The information processing device according to claim 1, wherein the specific surface is a floor surface of the virtual space, and the intersection points correspond to positions where each of the at least some of the people is standing on the floor surface.

8. The information processing device according to claim 1, wherein the virtual space has a 3D model corresponding to a non-architectural structure provided in the real space, the specific surface is a surface of the non-architectural structure exposed to the aisle, and the intersection is a position on the surface corresponding to a position touched by each hand of at least some of the people.

9. The information processing device according to claim 1, wherein the virtual space has a 3D model corresponding to a non-architectural structure provided in the real space, the specific surface is a surface of the non-architectural structure that is exposed to an aisle, and the intersection point corresponds to a position on the surface where the lines of sight of each of the at least some of the people intersect.

10. An information processing method comprising the steps of: acquiring video data showing a captured image of a real space; acquiring 3D model data of a virtual space showing said real space; and outputting behavioral data obtained by analyzing, for at least some of said plurality of people, the trajectories of intersections between a line obtained based on points included in 3D skeletal key points extracted from each of said plurality of people shown in said video and a specific surface of said virtual space.

Citation Information

Patent Citations

  • Information processing device, display control method, and program

    JP2019079298A

  • Object tracking device and program thereof

    JP2020107071A

  • Information processing device and marketing activity support device

    JP2021026336A

  • Computer-implemented method, data processing apparatus and computer program for generating three-dimensional pose-estimation data

    JP2022123843A

  • Camera localization based on skeletal tracking

    US20200273200A1