Image analysis apparatus, image analysis method, and image analysis program
Patent Information
- Application Number
- PCT/JP2025/022977
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2025-06-26
- Publication Date
- 2026-10-01
Smart Images

Figure JP2025022977_01102026_PF_FP_ABST
Abstract
Description
Image analysis apparatus, image analysis method, and image analysis program
[0001] The present disclosure relates to an image analysis apparatus, an image analysis method, and an image analysis program.
[0002] In the railway industry, the work status of station staff can only be grasped at coarse granularity after the fact based on work diaries, and there is a problem that optimal staffing cannot be achieved. In addition, developing technologies and solutions that realize mechanisms for maintaining business operations without degrading service quality even in a future where the number of station staff will decrease along with the decline in the working-age population has become an issue.
[0003] Patent Document 1 discloses a technology that uses a fixed camera installed near a work area to acquire image data of a worker (analysis subject), calculates a work value including a physical load value from the analysis subject's own behavior based on the acquired image data, and displays the calculated work value.
[0004] Japanese Patent No. 7487057
[0005] However, the technology disclosed in Patent Document 1 analyzes images acquired by a fixed camera installed near the work area. Since station staff all wear the same uniform, if two or more station staff are captured in the same camera, the station staff cannot be distinguished from each other. In order to distinguish individual station staff, it is necessary to use other sensors such as beacons, which poses a problem of enormous installation costs.
[0006] It is also conceivable to analyze work content from first-person perspective image data (hereinafter referred to as "first-person image") acquired by a wearable camera worn by a station staff who is an analysis subject, instead of using a fixed camera. However, when the technology disclosed in Patent Document 1 is applied to a first-person image, there is a problem that the work of the analysis subject cannot be estimated because the analysis subject wearing the wearable camera does not appear in the image.
[0007] An object of the present disclosure is to provide an image analysis apparatus, an image analysis method, and an image analysis program that can estimate the work content of an analysis subject from a first-person image captured by a wearable camera attached to the analysis subject.
[0008] The image analysis device disclosed herein comprises: a surrounding person behavior estimation unit that estimates the number of surrounding people present around the analysis subject in the input image and the state of those surrounding people based on a human skeleton extraction model extracted from an input image from a wearable camera that acquires image data at a field of view corresponding to the human field of view centered on the front of the analysis subject; a storage device that stores a work correspondence table showing the relationship between the number of surrounding people, the state of those surrounding people, and the work of the analysis subject; and a work estimation unit that refers to the work correspondence table and estimates and outputs the work of the analysis subject corresponding to the estimated number of surrounding people and the state of those surrounding people.
[0009] The image analysis method disclosed herein is an image analysis method performed by a computer, comprising the steps of: estimating the number of people surrounding the subject of analysis in the input image and the state of those people, based on a human skeleton extraction model extracted from an input image from a wearable camera that acquires image data at a field of view corresponding to the human field of view centered on the front of the subject of analysis; and estimating and outputting the work of the subject of analysis corresponding to the estimated number of people surrounding the subject and the state of those people, based on a work correspondence table that shows the relationship between the number of people surrounding the subject, the state of those people, and the work of the subject of analysis.
[0010] The image analysis program disclosed herein causes a computer to perform the following steps: estimate the number of people surrounding the subject of analysis in the input image and the state of those people, based on a human skeleton extraction model extracted from an input image from a wearable camera that acquires image data at a field of view corresponding to the human field of view centered on the front of the wearer, the subject of analysis; and estimate and output the work of the subject of analysis corresponding to the estimated number of people and the state of those people, based on a work correspondence table that shows the relationship between the number of people surrounding, the state of those people, and the work of the subject of analysis.
[0011] According to this disclosure, it is possible to provide an image analysis device, an image analysis method, and an image analysis program that can estimate the work content of a subject from first-person images taken by a wearable camera attached to the subject.
[0012] This is a block diagram showing an example of the hardware configuration of the image analysis device according to Embodiment 1. This is an example of a functional block diagram of the image analysis device according to Embodiment 1. This is a schematic diagram showing an example of a work correspondence table according to Embodiment 1. This is a flowchart showing an example of the processing of the image analysis device according to Embodiment 1. This is a block diagram showing an example of the hardware configuration of the image analysis device according to Embodiment 2. This is an example of a functional block diagram of the image analysis device according to Embodiment 2. This is an example of a functional block diagram of the work area estimation unit. This is a schematic diagram showing an example of a work correspondence table according to Embodiment 2. This is a flowchart showing an example of the processing of the image analysis device according to Embodiment 2. This is a block diagram showing an example of the hardware configuration of the image analysis device according to Embodiment 3. This is an example of a functional block diagram of the image analysis device according to Embodiment 3. This is an example of a functional block diagram of the language information generation unit related to the work area. This is a flowchart showing an example of the processing of the image analysis device according to Embodiment 3. This is a block diagram showing an example of the hardware configuration of the image analysis device according to Embodiment 4. This is an example of a functional block diagram of the image analysis device according to Embodiment 4. This is a schematic diagram showing an example of a work correspondence table according to Embodiment 4. This is a flowchart showing an example of the processing of the image analysis device according to Embodiment 4.
[0013] The image analysis apparatus, image analysis method, and image analysis program related to this disclosure will be described below with reference to the drawings. The following embodiments are merely examples, and it is possible to combine the embodiments as appropriate and modify each embodiment as appropriate.
[0014] Embodiment 1 Figure 1 is a block diagram showing an example of the hardware configuration of an image analysis device 10 according to Embodiment 1. As shown in Figure 1, the image analysis device 10 is composed of a computer in which a processor 20, a storage device 30, and an input / output interface 40, each of which are computing elements, are connected to a system bus 11. The image analysis device 10 may be composed of multiple computers connected by a network, or it may be composed of processing circuits.
[0015] The processor 20 is an IC (Integrated Circuit) that performs arithmetic processing such as a CPU (Central Processing Unit). In addition to a CPU, other arithmetic elements such as a DSP (Digital Signal Processor), GPU (Graphics Processing Unit), Network Processor, or FPGA (Field Programmable Gate Array) may also be used. By executing the image analysis program according to Embodiment 1, the processor 20 can realize an image analysis method that has a surrounding person behavior estimation function that estimates the number of people and the state of the surrounding people present around the person being analyzed, and a work estimation function that estimates the work of the person being analyzed wearing a wearable camera. As a result, by executing the image analysis program, the processor 20 functions as a surrounding person behavior estimation unit 21 and a work estimation unit 22. Furthermore, the image analysis program is provided, for example, on a recording medium on which these are recorded.
[0016] The storage device 30 consists of a volatile storage device such as RAM (Random Access Memory), a non-volatile storage device such as ROM (Read Only Memory), and a non-volatile storage device such as an HDD (Hard Disk Drive) or flash memory. Of the storage devices 30, the storage device stores various programs such as the OS (Operating System) and image analysis programs, as well as various data used to execute the image analysis programs. As an example of the various data, a work correspondence table 31 showing the relationship between the number of people in the surrounding area, the state of the people in the surrounding area, and the work of the person being analyzed is stored. The ROM stores computer startup programs such as the BIOS (Basic Input / Output System) in advance. The RAM is loaded with programs and data used for the operation of the processor 20 when the processor 20 is operating.
[0017] The input / output interface 40 is a port to which wired or wireless networks, input devices such as keyboards, mice, or touch panels, and output devices such as displays or printers are connected.
[0018] Figure 2 is an example of a functional block diagram of the image analysis device 10 according to Embodiment 1. As shown in Figure 2, the input image acquired by the wearable camera is input to the surrounding person behavior estimation unit 21 via the input / output interface 40. The surrounding person behavior estimation unit 21 uses a skeleton extraction model to extract the skeletons captured in the input image. The surrounding person behavior estimation unit 21 estimates and outputs the number of people surrounding the person being analyzed in the input image and their state (facing forward, sitting, or fighting) from the skeletons extracted from the input image. For example, OpenPose can be used as the skeleton extraction model. OpenPose is an open-source library that can estimate human posture from image data with high accuracy. The wearable camera is configured to acquire image data at a field of view corresponding to the human field of view centered on the front of the person being analyzed (for example, 120 to 130 degrees vertically and 150 to 200 degrees horizontally).
[0019] The task estimation unit 22 receives the number of people in the surrounding area and their status from the surrounding person behavior estimation unit 21. The task estimation unit 22 refers to the task correspondence table 31 and extracts the tasks of the person being analyzed that correspond to the number of people in the surrounding area and their status, thereby estimating and outputting the tasks of the person being analyzed who is wearing a wearable camera. In addition to using the task correspondence table 31, a neural network trained using a dataset in which the tasks of the person being analyzed are assigned as ground truth labels to the number of people in the surrounding area and their status may also be used to estimate the tasks of the person being analyzed. By using such a neural network, the number of people in the surrounding area and their status can be input into the neural network, and the tasks of the person being analyzed can be obtained as output.
[0020] The method for detecting a person's state from the results of a skeletal extraction model is as follows. For example, the state in which surrounding people are in front of the person being analyzed is when the following two conditions are met: (1) The number of detected people is 1. (2) The ratio (α / β) of the length of both shoulders (α) to the length of the arms (β) of the detected person is greater than or equal to a certain value (δ).
[0021] The state in which surrounding people are sitting down is determined when the following two conditions are met: (1) The number of detected people is 1. (2) The difference between the average Y coordinates of both shoulders and the average Y coordinates of both knees of the detected person is less than or equal to a certain value (ε).
[0022] A situation where people in the vicinity are fighting is considered to be met if the following three conditions are met: (1) The number of detected people is 2. (2) The distance between the keypoint (coordinate position) of one wrist and any keypoint on the other wrist, elbow, or shoulder is less than or equal to a certain value (η). (3) The movement speed of the wrist (distance from the wrist coordinate in the previous frame) is greater than or equal to a certain value (ω).
[0023] Each of the fixed values δ, ε, η, and ω used as the criteria for evaluation will be determined individually and specifically through prior empirical experiments, etc.
[0024] Figure 3 is a schematic diagram showing an example of a work correspondence table 31 according to Embodiment 1. As shown in Figure 3, the work correspondence table 31 shows the relationship between the number of people in the surrounding area, the state of those people, and the work of the person being analyzed.
[0025] For example, if there is one person in the surrounding area, and that person is facing the person being analyzed, then the person being analyzed is assisting passengers.
[0026] If there is one person in the surrounding area and that person is sitting down, the task of the person being analyzed is to respond to a medical emergency.
[0027] If there are two people around and they are fighting, the task of the person being analyzed is to mediate the fight.
[0028] If the number of people in the surrounding area is 0 and the status of the surrounding people cannot be detected, the person being analyzed will be on standby.
[0029] Figure 4 is a flowchart showing an example of the processing of the image analysis device 10 according to Embodiment 1. In step S101, the image analysis device 10 acquires an input image via the input / output interface 40.
[0030] In step S102, the surrounding person behavior estimation unit 21 estimates the number of people in the surrounding area and their status from the acquired image.
[0031] In step S103, the work estimation unit 22 estimates the work of the person being analyzed, who is wearing a wearable camera, based on the number of people in the surrounding area and their status, and then terminates the process.
[0032] As described above, the image analysis device 10 according to Embodiment 1 has a surrounding person behavior estimation unit 21 that estimates the number of people in the surrounding area and their status from the skeletons extracted from the input image, and a work estimation unit 22 that estimates the work of the person being analyzed based on the number of people in the surrounding area and their status, and the work correspondence table 31. As a result, it is possible to provide an image analysis device, an image analysis method, and an image analysis program that can estimate the work content of the person being analyzed from a first-person image taken by a wearable camera attached to the person being analyzed.
[0033] Embodiment 2 will now be described. Figure 5 is a block diagram showing an example of the hardware configuration of the image analysis device 12 according to Embodiment 2. The image analysis device 12 according to Embodiment 2 differs from the image analysis device 10 according to Embodiment 1 in that the storage device 30 stores language information 32 related to the work area of the person being analyzed and a work correspondence table 33 that takes into account the work area of the person being analyzed, and includes a work area estimation unit 23 that estimates the work area related to the language information with a high degree of similarity to the input image as the area where the person being analyzed is located. However, the other configurations are the same as in Embodiment 1, so the same reference numerals as in Embodiment 1 are used for the same configurations as in Embodiment 1, and detailed explanations are omitted.
[0034] Figure 6 is an example of a functional block diagram of the image analysis device 12 according to Embodiment 2. As shown in Figure 6, the image analysis device 12 according to Embodiment 2 receives input images via the input / output interface 40, similar to the image analysis device 10 according to Embodiment 1. The surrounding person behavior estimation unit 21 estimates and outputs the number of people in the surrounding area and their state (facing forward, sitting, or fighting) from the input image, just as in Embodiment 1. The input image is also input to the work area estimation unit 23 in addition to the surrounding person behavior estimation unit 21.
[0035] The work area estimation unit 23 vectorizes (extracts features from) the language information 32 related to the pre-set work area and the input image. The work area estimation unit 23 calculates the cosine similarity from each vector and estimates the work area related to the language information with high similarity to the image as the area where the person being analyzed is located.
[0036] The work estimation unit 24 takes the number of people in the surrounding area, the state of the people in the surrounding area, and the estimated work area as input, refers to the work correspondence table 33, and extracts the work of the person being analyzed that corresponds to the number of people in the surrounding area, the state of the people in the surrounding area, and the work area, thereby estimating and outputting the work of the person being analyzed wearing a wearable camera. For estimation, the work correspondence table 33 may be used, or a neural network trained using a dataset in which the person being analyzed's work is assigned as a ground truth label to the number of people in the surrounding area, the state of the people in the surrounding area, and the estimated work area may be used. By using such a neural network, the number of people in the surrounding area, the state of the people in the surrounding area, and the estimated work area are input to the neural network, and the person being analyzed's work is obtained as output.
[0037] Figure 7 is an example of a functional block diagram of the work area estimation unit 23. The work area estimation unit 23 includes an image feature extraction unit 50 that extracts features from an input image by vectorizing it, a language feature extraction unit 51 that extracts features from language information 32 related to the work area by vectorizing it, a similarity calculation unit 52 that calculates the cosine similarity of the image and language vectors, and a result determination unit 53 that outputs the work area associated with the language information with the highest cosine similarity as the estimation result.
[0038] The image feature extraction unit 50 takes an image as input and vectorizes it using a multimodal AI (Artificial Intelligence) model that can handle images and language in the same feature space. One example of a model used for vectorization is CLIP (Contrastive Language Image Pretraining). CLIP is a large-scale foundational model that can associate a series of images and text by jointly training an image encoder and a text encoder, and can handle language and images in the same feature space.
[0039] The language feature extraction unit 51 receives language information preset for each area as input, and vectorizes the language information using a model such as CLIP that can handle images and language in the same feature space. Vectorization of language information, for example, represents natural language as a set of a plurality of numerical values, and extracts this set of a plurality of numerical values as vector components.
[0040] The similarity calculation unit 52 calculates cosine similarity using the vectorized image and the vectorized language information. If the same model such as CLIP is used for both the image and the language information, the components of the vectorized image and the components of the vectorized language information for the same object will be approximate, so the calculated cosine similarity generally exhibits a high value.
[0041] The result determination unit 53 outputs, as an estimation result, the work area associated with the language information having the highest cosine similarity.
[0042] FIG. 8 is a schematic diagram showing an example of the work correspondence table 33 according to the second embodiment. As shown in FIG. 8, the work correspondence table 33 shows the relationship among the number of surrounding people, the state of surrounding people, work areas, and the work of the analysis target person.
[0043] For example, when the number of surrounding people is 1, the state of the surrounding people is in front of the analysis target person, and the work area is a counter, the work of the analysis target person is passenger handling (routine work).
[0044] When the number of surrounding people is 1, the state of the surrounding people is in front of the analysis target person, and the work area is a platform, the work of the analysis target person is passenger handling (non-routine work).
[0045] When the number of surrounding people is 2, the state of the surrounding people is fighting, and the work area is a platform, the work of the analysis target person is fight mediation (non-routine work).
[0046] When the number of surrounding people is 0, no surrounding people are detected, and the work area is a platform, the work of the analysis target person is train monitoring.
[0047] FIG. 9 is a flowchart showing an example of processing by the image analysis device 12 according to Embodiment 2. In FIG. 9, the same steps as those in Embodiment 1 are denoted by the same reference numerals as in Embodiment 1. In step S101, the image analysis device 12 acquires an input image via the input / output interface 40 in the same manner as in Embodiment 1.
[0048] In step S102, the surrounding person behavior estimation unit 21 estimates the number of surrounding persons and their states from the image acquired in the same manner as in Embodiment 1.
[0049] In step S201, the work area estimation unit 23 vectorizes each of the input image and the language information related to the work area using a multimodal model such as CLIP, and calculates the cosine similarity of the vectors, thereby extracting language information having high similarity to the input image and estimating the work area.
[0050] In step S202, the work estimation unit 24 estimates the work of the analysis subject wearing the wearable camera from the number of surrounding persons, their states, and the work area, and ends the process.
[0051] In the image analysis device 10 according to Embodiment 1, since the work of the analysis subject is estimated only from the number of surrounding persons and their behaviors, the types of work that can be estimated are limited. However, the image analysis device 12 according to Embodiment 2 includes, in addition to the configuration of Embodiment 1, a work area estimation unit 23 that estimates a work area from an input image captured by a wearable camera attached to the analysis subject. As a result, it is possible to estimate the work of the analysis subject in consideration of the work area of the analysis subject in addition to the number and behaviors of surrounding persons, so that the estimation accuracy can be improved. Furthermore, the work area estimation unit 23 can reduce the cost related to learning required for conventional image classification models by applying a multimodal AI model that can handle images and language in the same space to estimating the location of the work area.
[0052] Embodiment 3 will now be described. Figure 10 is a block diagram showing an example of the hardware configuration of the image analysis device 14 according to Embodiment 3. The image analysis device 14 according to Embodiment 3 differs from the image analysis device 12 according to Embodiment 2 in that it automates the selection of language information related to the work area. Specifically, it further includes a language information generation unit 25 related to the work area that extracts language information related to the work area from images related to the work area, and stores the generated language information 34 related to the work area in a storage device 30, which is different from the image analysis device 12 according to Embodiment 2. However, the other configurations are the same as in Embodiment 2, so the same reference numerals as in Embodiment 2 are used for the same configurations as in Embodiment 2, and detailed explanations are omitted.
[0053] Figure 11 is an example of a functional block diagram of the image analysis device 14 according to Embodiment 3. As shown in Figure 11, the image analysis device 14 according to Embodiment 3 receives input images via the input / output interface 40, similar to the image analysis device 10 according to Embodiment 1. The surrounding person behavior estimation unit 21 estimates and outputs the number of people in the surrounding area and their state (facing forward, sitting, or fighting) from the input image, just as in Embodiment 1. The input image is also input to the work area estimation unit 23 in addition to the surrounding person behavior estimation unit 21.
[0054] The language information generation unit 25 related to the work area receives multiple images related to the work area as input. The language information generation unit 25 converts each input image into language information using an image caption model. Furthermore, the language information generation unit 25 analyzes the converted language information and outputs, for example, the most frequently occurring word as the language related to that work area. The output language information 34 related to the work area is stored in the storage device 30.
[0055] The work area estimation unit 23 vectorizes (extracts features from) the language information 34 related to the pre-set work area and the input image, in the same manner as in Embodiment 2. The work area estimation unit 23 calculates the cosine similarity from each vector and estimates the work area related to the language information with high similarity to the image as the area where the person being analyzed is located.
[0056] The work estimation unit 24, similar to Embodiment 2, takes the number of people in the surrounding area, the state of the people in the surrounding area, and the estimated work area as input, refers to the work correspondence table 33, and extracts the work of the person being analyzed that corresponds to the number of people in the surrounding area, the state of the people in the surrounding area, and the work area, thereby estimating and outputting the work of the person being analyzed wearing a wearable camera. For estimation, the work correspondence table 33 may be used, or a neural network trained using a dataset in which the work of the person being analyzed is assigned as a ground truth label to the number of people in the surrounding area, the state of the people in the surrounding area, and the estimated work area may be used. By using such a neural network, the number of people in the surrounding area, the state of the people in the surrounding area, and the estimated work area are input to the neural network, and the work of the person being analyzed is obtained as output.
[0057] Figure 12 is an example of a functional block diagram of the language information generation unit 25 related to the work area. The language information generation unit 25 related to the work area includes a language information conversion unit 60 that takes an image related to the area as input and converts it into language information using an image caption model, and an extraction unit 61 that uses multiple converted language information to analyze the most frequent word and outputs the extracted most frequent word as the language related to the work area.
[0058] The image caption model used in the language information conversion unit 60 extracts image features using, for example, a trained CNN (Convolutional Neural Network). Then, the image caption model used in the language information conversion unit 60 extracts text features using, for example, LSTM (Long Short-Term Memory). Furthermore, the language information conversion unit 60 converts the content of the image into text by selecting text features that correspond to the extracted image features.
[0059] The extraction unit 61 outputs the most frequently occurring word from the multiple pieces of language information obtained by the language information conversion unit 60 as the language related to the work area.
[0060] Figure 13 is a flowchart showing an example of the processing of the image analysis device 14 according to Embodiment 3. In Figure 13, the same reference numerals are used for the same steps as in Embodiments 1 and 2. In step S301, the language information generation unit 25 related to the work area analyzes the images acquired for each area and generates language information related to the work area.
[0061] In step S101, the image analysis device 12 acquires an input image via the input / output interface 40, similar to embodiments 1 and 2.
[0062] In step S102, the surrounding person behavior estimation unit 21 estimates the number of people in the surrounding area and their status from the acquired images, similar to embodiments 1 and 2.
[0063] In step S201, the work area estimation unit 23 vectorizes the input image and the linguistic information related to the work area using a multimodal model such as CLIP, similar to Embodiment 2, and calculates the cosine similarity of the vectors to extract the linguistic information that has a high similarity to the input image and estimate the work area.
[0064] In step S202, the work estimation unit 24 estimates the number of people in the surrounding area, their status, and the work of the person being analyzed wearing a wearable camera, similar to Embodiment 2, and then terminates the process.
[0065] In the image analysis device 12 according to Embodiment 2, the linguistic information related to the work area, which is input to the work area estimation unit 23, must be selected manually, and this selection is labor-intensive. However, the image analysis device 14 according to Embodiment 3 inputs pre-acquired images related to the work area into an image caption model and obtains caption results. Furthermore, by selecting frequently occurring keywords from the caption results as the linguistic information related to that work area, the selection of linguistic information related to the work area can be automated, thereby reducing the labor involved in the selection process.
[0066] Embodiment 4 Next, Embodiment 4 will be described. Figure 14 is a block diagram showing an example of the hardware configuration of the image analysis device 16 according to Embodiment 4. The image analysis device 14 according to Embodiment 4 differs from the image analysis device 12 according to Embodiment 2 in that it adds voice as input and estimates the work of the person being analyzed from the conversation content, the number of people in the surroundings, the actions of the people in the surroundings, and the work area. Specifically, it further includes a conversation content estimation unit 27 that takes voice as input, converts the conversation content into text by speech recognition, and estimates the conversation content from the result, and differs from the image analysis device 12 according to Embodiment 2 in that it uses a work correspondence table 35 that shows the relationship between the number of people in the surroundings, the state of the people in the surroundings, the work area, the conversation content, and the work of the person being analyzed. However, the other configurations are the same as in Embodiment 2, so the same reference numerals as in Embodiment 2 are used for the same configurations as in Embodiment 2 and detailed explanations are omitted.
[0067] Figure 15 is an example of a functional block diagram of the image analysis device 16 according to Embodiment 4. As shown in Figure 15, the image analysis device 16 according to Embodiment 4 receives input images via the input / output interface 40, similar to the image analysis device 12 according to Embodiment 2. The surrounding person behavior estimation unit 21 estimates and outputs the number of people in the surrounding area and their state (facing forward, sitting, or fighting) from the input image, just as in Embodiment 2. The input image is also input to the work area estimation unit 23 in addition to the surrounding person behavior estimation unit 21.
[0068] The work area estimation unit 23, similar to Embodiment 2, vectorizes (extracts features from) the language information 32 related to the pre-set work area and the input image. The work area estimation unit 23, similar to Embodiment 2, calculates the cosine similarity from each vector and estimates the work area related to the language information with high similarity to the image as the area where the person being analyzed is located.
[0069] The conversation content estimation unit 27 receives the input audio. The input audio is acquired, for example, by a wearable camera worn by the person being analyzed. The conversation content estimation unit 27 converts the input audio into text using speech recognition and estimates the conversation content from the result. The estimation method may involve keyword matching, or the speech recognition result may be input into an AI such as chatGPT (registered trademark) to estimate the conversation content.
[0070] The work estimation unit 26 takes the number of people in the surrounding area, their status, estimated work area, and conversation content as input, refers to the work correspondence table 35, and extracts the work of the person being analyzed that corresponds to the number of people in the surrounding area, their status, conversation content, and work area, thereby estimating and outputting the work of the person being analyzed wearing a wearable camera. For estimation, the work correspondence table 35 may be used, or a neural network trained using a dataset in which the person being analyzed's work is assigned as a ground truth label to the number of people in the surrounding area, their status, estimated work area, and conversation content may be used. By using such a neural network, the number of people in the surrounding area, their status, estimated work area, and conversation content are input to the neural network, and the person being analyzed's work is obtained as output.
[0071] Figure 16 is a schematic diagram showing an example of a work correspondence table 35 according to Embodiment 4. As shown in Figure 16, the work correspondence table 35 shows the relationship between the number of people in the surrounding area, the state of the people in the surrounding area, the work area, the content of the conversation, and the work of the person being analyzed.
[0072] For example, if there is one person in the surrounding area, the person in the surrounding area is directly in front of the person being analyzed, the work area is a counter, and the conversation is about settlement, then the person being analyzed is performing settlement (routine work).
[0073] For example, if there is one person in the surrounding area, the person is directly in front of the person being analyzed, the work area is a counter, and the conversation is about giving directions, then the person being analyzed is giving directions (a routine task).
[0074] For example, if there is one person in the surrounding area, the person is directly in front of the person being analyzed, the work area is the home area, and the conversation is about giving directions, then the person being analyzed is giving directions (a non-routine task).
[0075] If there are two people around, they are fighting, and their conversation is argumentative, then the task of the person being analyzed is to mediate the fight.
[0076] If the number of people in the surrounding area is 0, the status of the surrounding people is not detected, and the work area is a platform, then the person being analyzed is performing train monitoring.
[0077] Figure 17 is a flowchart showing an example of the processing of the image analysis device 16 according to Embodiment 4.
[0078] In Figure 17, the same procedures as in Embodiments 1 and 2 are denoted by the same reference numerals as in Embodiments 1 and 2. In step S101, the image analysis device 12 acquires an input image via the input / output interface 40, similar to Embodiment 1.
[0079] In step S102, the surrounding person behavior estimation unit 21 estimates the number of people in the surrounding area and their status from the acquired images, similar to the first embodiment.
[0080] In step S201, the work area estimation unit 23, similar to Embodiment 2, vectorizes the input image and the linguistic information related to the work area using a multimodal model such as CLIP, calculates the cosine similarity of the vectors, extracts linguistic information with a high similarity to the input image, and estimates the work area.
[0081] In step S401, the conversation content estimation unit 27 takes, for example, audio acquired by a wearable camera as input, converts the conversation content into text using speech recognition, and estimates the conversation content from the result.
[0082] In step S402, the work estimation unit 26 estimates the work of the person being analyzed, who is wearing a wearable camera, based on the number of people in the surrounding area, their condition, the work area, and the content of their conversations, and then terminates the process.
[0083] As described above, the image analysis device 16 according to Embodiment 4 adds audio as input to the configuration of Embodiment 2, and estimates the work of the person being analyzed from the conversation content, the number of people in the surroundings, the state of the people in the surroundings, and the work area. In Embodiment 2, work estimation was limited because it was performed using only images, but according to Embodiment 4, the work content is estimated from the conversation content in addition to the number of people in the surroundings, the actions of the people in the surroundings, and the area where the person being analyzed is located, so the estimation accuracy can be improved.
[0084] 10, 12, 14, 16 Image analysis device, 20 Processor, 21 Surrounding person behavior estimation unit, 22 Task estimation unit, 23 Work area estimation unit, 24 Task estimation unit, 25 Language information generation unit related to the work area, 26 Task estimation unit, 27 Conversation content estimation unit, 30 Storage device, 31, 33, 35 Task correspondence table, 32, 34 Language information related to the work area, 40 Input / output interface, 50 Image feature extraction unit, 51 Language feature extraction unit, 52 Similarity calculation unit, 53 Result determination unit, 60 Language information conversion unit, 61 Extraction unit.
Claims
1. An image analysis device comprising: a surrounding person behavior estimation unit that estimates the number of surrounding people and the state of the surrounding people present around the analysis subject in the input image, based on a human skeleton extraction model extracted from input images from a wearable camera that acquires image data at a field of view corresponding to the human field of view centered on the front of the analysis subject; a storage device that stores a work correspondence table showing the relationship between the number of surrounding people, the state of the surrounding people, and the work of the analysis subject; and a work estimation unit that refers to the work correspondence table and estimates and outputs the work of the analysis subject corresponding to the estimated number of surrounding people and the state of the surrounding people.
2. The image analysis device according to claim 1, further comprising a work area estimation unit that calculates the similarity between the input image and language information related to the work area stored in the storage device, and estimates the work area related to the language information with a high similarity as the area where the person to be analyzed is located, wherein the work correspondence table shows the relationship between the number of people in the surrounding area, the state of the people in the surrounding area, the work area, and the work of the person to be analyzed, and the work estimation unit refers to the work correspondence table and estimates and outputs the work of the person to be analyzed corresponding to the estimated number of people in the surrounding area, the state of the people in the surrounding area, and the work area.
3. The image analysis device according to claim 2, wherein the work area estimation unit vectorizes the input image and the language information relating to the work area, and outputs as an estimation result the work area associated with the language information relating to the work area that has the highest cosine similarity calculated between the vectorized input image and the vectorized image.
4. The image analysis device according to claim 2 or 3, further comprising a work area language information generation unit that extracts language information related to the work area from an input image of the work area.
5. An image analysis device according to any one of claims 2 to 4, further comprising a conversation content estimation unit that recognizes audio acquired by the wearable camera and estimates the content of a conversation contained in the audio, wherein the work correspondence table shows the relationship between the number of people in the surrounding area, the state of the people in the surrounding area, the work area, the content of the conversation, and the work of the person being analyzed, and the work estimation unit refers to the work correspondence table and estimates and outputs the work of the person being analyzed that corresponds to the estimated number of people in the surrounding area, the state of the people in the surrounding area, the content of the conversation, and the work area.
6. An image analysis method performed by a computer, comprising the steps of: estimating the number of people surrounding the subject of analysis in the input image and the state of those people, based on a human skeleton extraction model extracted from an input image from a wearable camera that acquires image data at a field of view corresponding to the human field of view centered on the front of the subject of analysis; and estimating and outputting the work of the subject of analysis corresponding to the estimated number of people surrounding the subject and the state of those people, based on a work correspondence table showing the relationship between the number of people surrounding the subject, the state of those people, and the work of the subject of analysis.
7. An image analysis program that causes a computer to perform the following steps:
1. Estimate the number of people surrounding the subject of analysis in the input image and the state of those people, based on a human skeleton extraction model extracted from an input image from a wearable camera that acquires image data at a field of view corresponding to the human field of view centered on the front of the wearer, the subject of analysis; and 2. Estimate and output the work of the subject of analysis corresponding to the estimated number of people surrounding the subject and the state of those people, based on a work correspondence table showing the relationship between the number of people surrounding the subject, the state of those people, and the work of the subject of analysis.